The cure for AI slop is a 1986 aircraft manual/kit

experiment-results-openai.md

1 KB·Raw on GitHub ↗

Per-category detail for the gpt-5.5 run

STE experiment results — OpenAI (gpt-5.5)

6 prompts x 4 conditions. Violations per 100 words (lower = cleaner). em-dash = true dash chars only.

Conditionwordsviolationsper 100wvs baselineem-dashes
baseline649233.540
banwords608132.14-40%0
orwell53291.69-52%0
ste45481.76-50%0

Per-category (per 100 words)#

categorybaselinebanwordsorwellste
long_sentence(>20w)0.770.490.380.0
semicolon0.00.00.00.0
contraction0.150.00.00.0
passive_voice1.390.990.750.66
ing_main_verb0.460.160.190.0
nominalization0.460.160.00.22
phrasal_verb0.00.00.00.0
banned_word0.00.00.00.0
marketing_adjective0.00.00.00.0
modal_hedge0.00.00.00.0
long_paragraph(>6s)0.310.330.380.88

Headline: STE cut violations by 50% vs baseline (3.54 -> 1.76 per 100 words). Ban-words 40%, Orwell 52%.

← Back to the episode