Cloning into '/tmp/simpleenglish-functional'... == repo's own unit tests == ---------------------------------------------------------------------- Ran 7 tests in 0.003s OK == linter self-test (built-in fixtures) == self-test OK: 12 violations in slop fixture, 0 in clean == fresh-input probes == -- slop text (README's real before example): violations_total: 7 | per_100w: 16.28 | longest sentence: 26 words -- clean STE rewrite: violations_total: 0 | per_100w: 0.0 | longest sentence: 12 words -- edge cases (code-span-as-one-word + trailing conditions): sentence_over_limit: 0 (9-flag command counted as 8 words) | trailing_condition: 2 == independent reproduction of the headline number == model base/100w skill/100w reduction claude-opus-4-5-20251101 2.55 0.57 77.6% claude-opus-4-6 2.24 0.40 82.3% claude-opus-4-7 2.28 0.42 81.7% claude-opus-4-8 1.05 0.62 41.3% claude-opus-5 2.13 0.32 85.2% claude-sonnet-4-6 2.06 0.52 74.9% claude-sonnet-5 2.67 0.53 80.0% raw files scanned: 168; models with both conditions: 7 mean reduction across models: 74.7% (README claims 74.6%)