v3.2 stress tests
100 synthetic profiles run through the real scoring model, plus monotonicity, confidence-integrity, and influence tests. Nothing is persisted. All weights remain labeled heuristics.
Synthetic-profile distribution
Monotonicity — stronger answer must not lower the score
51 transitions tested across all scored questions.
Confidence-integrity tests
Domain influence — practical swing vs intended weight
Swing = overall score movement when a domain goes from its weakest to strongest plausible state (other domains neutral). Ratio = swing / intended weight.
| Domain | Min | Max | Swing | Intended | Ratio |
|---|---|---|---|---|---|
| Identity & Access | 55 | 76 | 21 | 25 | 84% |
| Passwords & Accounts | 60 | 80 | 20 | 25 | 80% |
| Device Security | 57 | 73 | 16 | 20 | 80% |
| Privacy & Exposure | 65 | 75 | 10 | 15 | 67% |
| Recovery & Resilience | 65 | 76 | 11 | 15 | 73% |
One or more domains deviate from intended influence (flagged in amber).
Question influence — max practical impact
Max overall Security-Score swing from moving a single question from its weakest to strongest answer (others neutral). Flagged if > 12 points.
| Question | Domain | Sec swing | Conf swing | Flag |
|---|---|---|---|---|
| id_mfa | Identity | 10 | -1 | — |
| id_mfa_method | Identity | 5 | 1 | — |
| id_mfa_breadth | Identity | 5 | 0 | — |
| pw_manager | Passwords | 7 | 5 | — |
| pw_reuse | Passwords | 8 | 5 | — |
| pw_reuse_breadth | Passwords | 5 | 0 | — |
| dv_updates | Device | 5 | 0 | — |
| dv_encryption | Device | 6 | -2 | — |
| dv_downloads | Device | 4 | 0 | — |
| pr_accounts_public | Privacy | 4 | 0 | — |
| pr_overshare | Privacy | 3 | 0 | — |
| pr_permissions | Privacy | 3 | 0 | — |
| rc_backups | Recovery | 6 | 0 | — |
| rc_recovery | Recovery | 4 | 0 | — |
No single question owns more than 60% of its domain.
Adversarial personas
Deliberately picks every answer that maximizes the score, with full certainty and self-confirmation.
Expected: Security 100. Confidence high but capped below 100 (self-reported ceiling).
Claims strong controls with full certainty, but provides conflicting follow-ups.
Expected: High security, but confidence reduced by two conflicts (MFA method + password reuse).
2 conflict(s) detected.
Strong actual behaviors but many uncertain responses and no self-confirmation.
Expected: High-ish security, moderate confidence (uncertain + unconfirmed).
Clearly understands their poor controls and answers precisely and consistently.
Expected: Low security, but high confidence (honest, specific, consistent).
Strong everywhere except no MFA on the primary email.
Expected: Security well below 100 (identity domain tanks) but still strong elsewhere; confidence high.