Script AI · Retention Science

How Script AI Retention Scoring Works: Predicting Performance Before You Record

AI retention scoring analyzes 47 structural and linguistic features in your script and predicts where viewers will drop off — with a 0.74 correlation to actual audience behavior. That's not a guess. That's a structural audit. The model has been calibrated against 1,200 published scripts with known retention data. It does not measure charisma. It does not measure topic appeal. It measures whether your structure keeps attention or bleeds it — and that turns out to be the part you can control before pressing record.

What the Score Actually Predicts (and What It Doesn't)

A 0.74 correlation is high for behavioral prediction. For context: credit scores correlate with loan repayment at 0.35-0.45. College admissions tests correlate with first-year GPA at 0.35-0.50. The model does not predict view counts. It predicts retention — the percentage of viewers still watching at specific timestamps. The relationship between retention and views runs through the recommendation engine, and that relationship is noisy. A retention score of 78 predicts the structure will hold 72%+ of viewers past 30 seconds. It does not predict whether YouTube will surface the video to enough people to matter.

Score RangePredicted 30s RetentionPredicted 5min RetentionActual Accuracy
85-10078%+58%+±7.2pp
70-8463-77%41-57%±9.1pp
55-6947-62%28-40%±11.4pp
Below 55Below 47%Below 28%±14.8pp

Accuracy degrades at the bottom because bad scripts fail for reasons the model can see — and for reasons it cannot. A script scoring 47 is definitely going to lose viewers. Whether it loses 38% or 56% depends on editing, charisma, and topic urgency. The model gets less precise the worse the script gets, but the practical takeaway is the same either way: rewrite it.

The 47-Feature Scoring Architecture

The model evaluates scripts across five categories, each contributing to the composite score with different weightings. Hook architecture carries the most weight (31% of total score) because the first 3 seconds explain more retention variance than any other single factor. Structural integrity follows at 27%. Pacing structure at 18%. Language complexity at 14%. Engagement triggers at 10%. These weights are not opinion — they are derived from regression analysis of which feature categories most accurately predict retention outcomes across the 1,200-script calibration set.

Hook Architecture (31% weight)

Six features: pattern-interrupt type classification, curiosity gap size measured in information units, promise specificity score, hook-to-title alignment ratio, first-sentence reading level, and hook density (hook words per second in first 5 seconds). A script with a Contrarian Data Point hook averaging 14 words in the first 4 seconds scores 12 points higher on hook architecture than a Welcome Back opener with 31 words in the same window.

Structural Integrity (27% weight)

Eighteen features including act balance ratio, payoff density distribution, foreshadowing-to-resolution interval, transition quality scoring, and the critical midpoint-sag predictor. Scripts where the gap between payoffs exceeds 90 seconds in the 2-5 minute window score 6-8 points lower on structural integrity than scripts maintaining 45-75 second payoff intervals.

Pacing Structure (18% weight)

Eight features: pattern-interrupt frequency, beat-length variance, information density per minute, section transition speed, emotional arc slope, sentence-length rhythm ratio, question-to-statement ratio, and the critical 45-second attention-reset compliance score. Scripts that miss two or more consecutive 45-second attention-reset windows lose an average of 14% more viewers than compliant scripts.

Language Complexity (14% weight)

Nine features: Flesch-Kincaid grade level, sentence-length variance, passive-voice ratio, jargon density, adverb frequency, filler-word count, concrete-vs-abstract word ratio, repetition score, and pronoun density. Scripts at a 7th-8th grade reading level with sentence-length variance above 9 words outperform college-level scripts with low variance by 18-22 points on retention.

Engagement Triggers (10% weight)

Six features: open-loop count and distribution, rhetorical question density, direct-audience-address frequency, emotional peak count, surprise-element scoring, and CTA placement optimization. Scripts with 3-5 open loops distributed across acts retain 28% more viewers past minute 5 than scripts with zero or more than 8 open loops.

Where Scoring Goes Wrong: The 37% Error Cases

63% of predictions land within 10 percentage points of actual retention. The 37% that miss cluster into three patterns. Each is instructive about what the model cannot see — and what creators should not blame on script quality when the real problem sits elsewhere.

Creator Charisma (41% of errors)

A script scoring 82 on structure can underperform by 18 points if the on-camera delivery is flat, monotone, or nervous. Conversely, a 61-scoring script can overperform by 14 points with exceptional delivery. The model scores words. Viewers hear words delivered by a person. That gap accounts for the largest error cluster.

Editing Quality (33% of errors)

A script cannot encode jump cuts, B-roll quality, background music, or pacing reinforcement through editing. A structurally sound script with a 0.8-second average cut rate underperforms a structurally sound script with a 0.4-second average cut rate by an estimated 8-12 retention points in fast-paced niches like tech and gaming.

Topic Novelty (26% of errors)

First-mover content on emerging topics retains differently than established-topic content. Viewers tolerate weaker structure when the information is genuinely new. The model was trained on scripts across both types but cannot determine topic novelty from text alone. A script about a breaking news event in a slow niche can outperform its score by 15+ points simply because viewers are information-hungry.

Scoring vs. General AI: The 0.11 Correlation Problem

When we submitted the same 200 scripts to ChatGPT and the Astryx scoring model, ChatGPT's quality evaluations showed a 0.11 correlation with actual retention. The model showed 0.74. The gap is not about intelligence. ChatGPT is arguably a more sophisticated system. The gap is about training data. ChatGPT was trained on all human-written text, where quality correlates with grammar, coherence, and factual accuracy — none of which predict whether someone keeps watching a video.

The scoring model was trained on scripts paired with retention data. It learned that short, punchy sentences correlate with higher retention than grammatically elegant ones. It learned that pattern interrupts placed at 47-second intervals outperform those placed at 90-second intervals. It learned that 7th-grade reading level scripts with high sentence-length variance outperform 12th-grade scripts with consistent structure. These are counterintuitive findings that general AI, trained on all text, cannot surface because its training data says good writing is consistent writing. Retention data says good YouTube writing is deliberately inconsistent writing. See how this differs from AI scripts vs human scripts — the patterns overlap but the mechanism is different.

The 3-Score Threshold: When to Publish, Revise, or Abandon

Creators who set explicit scoring thresholds publish better content with less indecision. Our data suggests three decision boundaries based on the composite score:

Composite ScoreDecisionExpected 30s RetentionAction
78+Publish-ready72%+Record. Only revise if sub-scores flag specific weak sections.
62-77Revise targeted sections52-71%Check which sub-score is lowest. Fix that section only. Rescore. Repeat max 2x.
Below 62Major restructure or abandonBelow 52%Structural problems exceed fixable scope. Restructure from outline or reconsider topic.

The most common mistake: creators scoring at 68 who spend 3 hours chasing an 80. The marginal gain from 68 to 72 is real but small — roughly 3-5 retention points. The marginal gain from 68 to a properly executed delivery and tight editing is larger — roughly 8-11 retention points. Know when the script is good enough and the bottleneck has shifted to execution. Our 3-pass editing process covers what to fix after the script is structurally sound.

What Creators Get Wrong About Scoring

Three misconceptions show up in nearly every creator conversation about retention scoring. They persist because they are intuitive — and wrong.

Misconception 1: "A high score means a viral video." The correlation between retention scores and view counts is 0.18. Structure does not create virality. Virality requires structure, topic timing, shareability mechanics, and algorithmic surface area. A 92-scoring script on a topic nobody searches will get 200 views. A 72-scoring script on a trending topic with a strong thumbnail can get 200,000. Score measures retention potential. Views require everything else.

Misconception 2: "I should score every draft and only publish above 80." Perfectionism kills channels faster than mediocre scripts. Creators who score every draft and refuse to publish below 80 publish 73% fewer videos than creators who use a 65 threshold. At 4x monthly output, the 65-threshold creator often has more total watch time than the 80-threshold creator — even at lower average retention — because volume compounds. Use scoring as a gate, not a prison.

Misconception 3: "Scoring replaces human judgment." The model flags structural problems. It does not flag boring topics, unlikeable delivery, or factually wrong claims. If your script scores 81 but the topic is a retread of something 14 other channels covered this week, the score is irrelevant. Retention scoring is a structural quality check. It is not a replacement for knowing your niche. For more on scoring versus manual judgment, see our guide to AI script A/B testing.

Next Steps

Want to know if your script will retain before you record?

Astryx scores your script against 47 retention features, predicts exactly where viewers will drop off, and tells you what to fix — before you spend hours filming and editing.

Try Astryx Free →