Build Log
Self-Consistency v2: Clean 21/21, Success Criterion Met
August 14, 2026
Following the exact next action named in the v1 negative result: a stricter adversarial prompt (reasonableness framing plus a rebuttal-question checkpoint), tested against the original 17 cases plus 4 new held-out clear-cut cases. Zero false escalations, both known ambiguities correctly caught, success criterion met on the first real run.
Timeline
What shipped
experiments/helm-classifier-probe/classify-with-self-consistency-v2.mjs — the fixed adversarial pass: reasonableness framing ('a technically-constructible argument is not enough') plus an explicit REBUTTAL_CHECK the model must answer before it's allowed to output CAN_ARGUE_DIFFERENT: YES
experiments/helm-classifier-probe/held-out-cases.json — 4 new, independently code-verified clear-cut test cases (all CONTENT-bucket — see the honest coverage-gap note below), not used to design the v2 prompt
experiments/helm-classifier-probe/results-self-consistency-v2.json — full real run output including every case's REBUTTAL_CHECK reasoning, not just the final bucket
“A real engineer who got this request would genuinely pause and ask 'does she want me to flip one event's status, or does she want a cap enforced going forward?' — these are materially different tasks and the request truly doesn't say.”