Rankings
The top 5 AI models for YouTube scriptwriting, by craft score. Judges graded every script on six writing skills; how each script fared against the real video it was based on is shown separately. The thin lines show the range each score could fall in. Opus 5 leads, but the group behind it is close.
LLM Leaderboard
Every LLM we benchmarked for scriptwriting: craft score, cost compared with the cheapest model, and writing time. Rank the board by a single skill to see which model writes the best hooks, storytelling, or voice match, and open a row for the full six-skill breakdown.
| Model | Overall | Cost | Speed | ||
|---|---|---|---|---|---|
| 1 | Claude Opus 5Anthropic | 87.9% | 3× | ~8.3 min | |
Best overall. It wrote a full script for all 12 tasks, earned the highest craft scores in the test, and was the most consistent model from task to task.
| |||||
| 2 | Claude Opus 4.8Anthropic | 84.4% | 2.3× | ~6.4 min | |
| 3 | XGrok 4.5xAI | 82.8% | 2.1× | ~6.8 min | |
| 4 | Claude Fable 5Anthropic | 82.7% | 5× | ~7.5 min | |
| 5 | K3Kimi K3Moonshot AI | 80.5% | 1× | ~10 min | |
| 6 | GPT-5.5OpenAI | 80.3% | 1.9× | ~4.7 min | |
| 7 | Claude Sonnet 5Anthropic | 79.3% | 1.2× | ~7.5 min | |
| 8 | MMuse Spark 1.1Meta | 79% | 1.8× | ~4.5 min | |
| 9 | GLMGLM 5.2Zhipu AI | 77.8% | 1× | ~14.6 min | |
| 10 | GPT-5.6 SolOpenAI | 75.5% | 1.7× | ~4.1 min | |
| 11 | Claude Sonnet 4.6Anthropic | 75.4% | 1.2× | ~7.3 min | |
Cost vs performance
Each AI model's craft score against the cost of one script. Higher up means better. Further left means cheaper.
Key takeaways
Which AI model to pick for YouTube scriptwriting, based on what you care about most: quality, reliability, cost, or speed.
Best overall: Claude Opus 5
A clear number one. It wrote a full script for all 12 tasks, got the highest craft scores from the judges, and was the most consistent model from task to task.
Best value: Kimi K3
Fifth overall at the lowest cost (tied with GLM 5.2). It is slower, at about 10 minutes per script.
Best writing, with a catch: Claude Fable 5
Craft scores tied with Opus 5, but its strict safety rules made it refuse two of the 12 scripts: a disturbing sea creatures topic and a CIA hacking story. A refused task gets the worst score of that round, which drops it to 4th.
Fastest: GPT-5.6 Sol
The quickest in the test at about 4 minutes per script, but near the bottom on craft, with the weakest voice match.
Surprising results
The findings we did not see coming.
GPT-5.6 Sol: great at code, weak at scripts
GPT-5.6 Sol is a top model on coding benchmarks, but it landed near the bottom here: below the older GPT-5.5, with a failed task and the weakest voice match in the test. Tuning a model that hard for code seems to come at a cost in creative work like scriptwriting.
Grok 4.5 is the dark horse
Nobody talks about Grok for YouTube scripts, yet it landed 3rd, ahead of GPT-5.5 and two of the four Claude models. It matched channel voices better than most, at a middle price.
Fable 5: Refused writing scripts
Fable 5 scored level with the winner and still fell to 4th, only because it refused two topics: a video about disturbing deep-sea creatures and one about the teens who talked their way into the CIA. A model you cannot count on loses points, no matter how well it writes.
Your next video starts here.
Write yours with the same tool. Give the AI Scriptwriter an idea and get back a full script, researched and written in your channel voice.
Craft comparison
How the top 5 AI scriptwriting models score on the six craft skills: hook, retention, storytelling, voice match, pacing, and clarity. Judges scored each from 1 to 10. The other models are in the leaderboard rows.
Generated scripts
Each task is based on a real hit video. Open one to read the full script each LLM wrote and see its scores. All scripts come from TubeLab's AI Scriptwriter.
How we tested it
How we benchmarked each LLM at YouTube scriptwriting, start to finish.
- 1
Real hit videos
We picked real hit YouTube videos from different genres. Each task asks the models to write that video again, and the real script serves as the reference point.
- 2
One tool, many models
Every model wrote inside the TubeLab AI Scriptwriter, the same tool creators use. Same voice training, same research, same steps. Only the model changed.
- 3
Same inputs for everyone
Every model got the same short idea and a voice trained on the original channel, then wrote a full script from scratch. Research was locked per task, so every model worked from the same facts.
- 4
Blind judging
Three AI judges from different companies (OpenAI, Anthropic, Google) grade every script on six writing skills, and models are ranked by those craft scores. They also grade the real video's script the same way, which shows how each script compares with the original. Judges never know which model wrote what, and no judge grades scripts from its own company. The finer scoring rules are in the FAQ.
FAQ
- How is the overall quality score calculated?
- Every script is graded on a six-skill rubric (hook, retention structure, storytelling, voice match, pacing, clarity), and a model's overall score is its average across those six skills and all tasks. Each task is scored by a panel of three judge models from different families (OpenAI, Anthropic, Google); judges never see which model wrote a script, and no judge scores scripts written by its own model family. Tasks a model refuses or fails are scored as the worst craft score any model achieved on that task, so a refusal can never inflate an average. The comparison against the original video is reported separately and does not drive the ranking.
- How do the scripts compare with the real videos?
- Each generated script and the original video's transcript get graded on the same rubric by the same judge. If the generated script scores more than a quarter point higher, it is rated above the original; more than a quarter point lower, below; within the band, it is on par. These comparisons are shown for context and do not affect the ranking, which comes from craft scores alone. Two caveats still apply: this measures what AI judges think of the text, not what audiences would watch, and a transcript strips out the real video's delivery, editing, thumbnail, and proven real-world performance.
- Why did some models not complete every task?
- Claude Fable 5 refused to write two scripts on content-safety grounds, and GPT-5.6 Sol failed to complete the writing flow on one task. These are penalized in the overall score, counted as the worst craft score any model achieved on that task: for a creator, a model that will not write the script has failed it. The craft score covers completed tasks only, and both numbers are shown.
- How are cost and speed measured?
- Cost is the average provider API cost of producing one finished script, including the voice-training and research steps. Speed is the average wall-clock time from brief to finished script.
- Can I write YouTube scripts with these AI models myself?
- Yes. The TubeLab AI Scriptwriter lets you pick the LLM that writes for you, including the top-ranked ones, with a voice trained on your own channel.
- How much do the judges agree with each other?
- On the script-vs-original verdicts, the panel was unanimous 46% of the time, with 52% mean pairwise agreement and a Fleiss' kappa of 0.24 over full three-judge cells (fair agreement beyond chance; many disagreements are win-vs-tie calls near the quarter-point band). Every overall score carries a bootstrap 95% confidence interval, shown on the rankings chart; models with overlapping intervals should be read as ties. The judges have not yet been validated against human ratings or real audience retention.
- Why not just ask the judges to pick the better script?
- We did at first, and the measurements showed it was biased. In forced head-to-head picks the judges never once called a tie, and our most neutral judge (Gemini, with no model in the race) preferred the AI script 76% of the time, more than either judge whose family competes. LLM judges are known to favor LLM-style prose, and the human side competes as a transcript stripped of delivery and visuals. Grading both scripts on the same rubric with a tie band fixed the worst of it: ties reappeared, and the weakest models now lose to the originals, which the old method had hidden.
- Does this pick the best model for scriptwriting in general?
- No. It picks the best model inside the TubeLab Scriptwriter, which is the choice this page is built to help with. Every model ran the same pipeline: the same prompts, the same tools, the same writing steps. A model that fits that pipeline well scores better in it, and a different tool with different prompts could order these models differently. What this measures is each model doing the full job, from research to finished script, not answering one prompt in a vacuum.
- Do different models get scored by different judges?
- Yes. No judge grades scripts from its own company, so Claude scripts are graded by the OpenAI and Google judges, GPT scripts by the Anthropic and Google judges, and models with no judge in the race by all three. We checked whether this skews the board. On the scripts all three judges graded, their average scores sit within about a quarter point of each other on the 10-point scale, and the Anthropic judge turned out to be the most lenient, which means the strictest panel is the one grading Claude models. We also re-ranked every model using only the Google judge, the one judge that scores everyone: the top four positions do not change. A judge panel with no competing families at all is planned for the next update.
- Who runs this benchmark, and what is the conflict of interest?
- TubeLab builds the Scriptwriter used to run every task, and this page promotes it. That conflict cannot be removed, so the design leans on verifiability instead: blind judging by three model families with a self-preference guard, refusals scored as the worst result on the task, published methodology and rubric, and confidence intervals on every score.
- How often is the benchmark updated?
- The benchmark was run in July 2026. We re-run it as significant new models ship; closely ranked models should be read as statistical ties.
Conclusion
What we would pick after judging every script.
Claude Opus 5 is the best AI model for YouTube scriptwriting in this test. It wrote a full script for all 12 tasks, posted the highest craft scores on the board, and judges rated it on par with the original video in most matchups.
The right pick still depends on what you care about. Kimi K3 delivers the most quality per dollar at the lowest price per script. Grok 4.5 balances price, speed, and craft. Claude Fable 5 writes at the same level as Opus 5, but its safety rules refused two topics, and a model that will not write can leave you without a script.
Keep the limits in mind: the judges are AI models, the sample is small, and close scores should be read as ties. Every script on this page came from the TubeLab AI Scriptwriter, so the fastest way to form your own opinion is to pick a model and write with it.
Your subscribers are waiting...
Type an idea. The AI Scriptwriter researches it, plans it, and writes the full script in your channel voice.
Rankings
The top 5 AI models for YouTube scriptwriting, by craft score. Judges graded every script on six writing skills; how each script fared against the real video it was based on is shown separately. The thin lines show the range each score could fall in. Opus 5 leads, but the group behind it is close.
LLM Leaderboard
Every LLM we benchmarked for scriptwriting: craft score, cost compared with the cheapest model, and writing time. Rank the board by a single skill to see which model writes the best hooks, storytelling, or voice match, and open a row for the full six-skill breakdown.
| Model | Overall | Cost | Speed | ||
|---|---|---|---|---|---|
| 1 | Claude Opus 5Anthropic | 87.9% | 3× | ~8.3 min | |
Best overall. It wrote a full script for all 12 tasks, earned the highest craft scores in the test, and was the most consistent model from task to task.
| |||||
| 2 | Claude Opus 4.8Anthropic | 84.4% | 2.3× | ~6.4 min | |
| 3 | XGrok 4.5xAI | 82.8% | 2.1× | ~6.8 min | |
| 4 | Claude Fable 5Anthropic | 82.7% | 5× | ~7.5 min | |
| 5 | K3Kimi K3Moonshot AI | 80.5% | 1× | ~10 min | |
| 6 | GPT-5.5OpenAI | 80.3% | 1.9× | ~4.7 min | |
| 7 | Claude Sonnet 5Anthropic | 79.3% | 1.2× | ~7.5 min | |
| 8 | MMuse Spark 1.1Meta | 79% | 1.8× | ~4.5 min | |
| 9 | GLMGLM 5.2Zhipu AI | 77.8% | 1× | ~14.6 min | |
| 10 | GPT-5.6 SolOpenAI | 75.5% | 1.7× | ~4.1 min | |
| 11 | Claude Sonnet 4.6Anthropic | 75.4% | 1.2× | ~7.3 min | |
Cost vs performance
Each AI model's craft score against the cost of one script. Higher up means better. Further left means cheaper.
Key takeaways
Which AI model to pick for YouTube scriptwriting, based on what you care about most: quality, reliability, cost, or speed.
Best overall: Claude Opus 5
A clear number one. It wrote a full script for all 12 tasks, got the highest craft scores from the judges, and was the most consistent model from task to task.
Best value: Kimi K3
Fifth overall at the lowest cost (tied with GLM 5.2). It is slower, at about 10 minutes per script.
Best writing, with a catch: Claude Fable 5
Craft scores tied with Opus 5, but its strict safety rules made it refuse two of the 12 scripts: a disturbing sea creatures topic and a CIA hacking story. A refused task gets the worst score of that round, which drops it to 4th.
Fastest: GPT-5.6 Sol
The quickest in the test at about 4 minutes per script, but near the bottom on craft, with the weakest voice match.
Surprising results
The findings we did not see coming.
GPT-5.6 Sol: great at code, weak at scripts
GPT-5.6 Sol is a top model on coding benchmarks, but it landed near the bottom here: below the older GPT-5.5, with a failed task and the weakest voice match in the test. Tuning a model that hard for code seems to come at a cost in creative work like scriptwriting.
Grok 4.5 is the dark horse
Nobody talks about Grok for YouTube scripts, yet it landed 3rd, ahead of GPT-5.5 and two of the four Claude models. It matched channel voices better than most, at a middle price.
Fable 5: Refused writing scripts
Fable 5 scored level with the winner and still fell to 4th, only because it refused two topics: a video about disturbing deep-sea creatures and one about the teens who talked their way into the CIA. A model you cannot count on loses points, no matter how well it writes.
Your next video starts here.
Write yours with the same tool. Give the AI Scriptwriter an idea and get back a full script, researched and written in your channel voice.
Craft comparison
How the top 5 AI scriptwriting models score on the six craft skills: hook, retention, storytelling, voice match, pacing, and clarity. Judges scored each from 1 to 10. The other models are in the leaderboard rows.
Generated scripts
Each task is based on a real hit video. Open one to read the full script each LLM wrote and see its scores. All scripts come from TubeLab's AI Scriptwriter.
How we tested it
How we benchmarked each LLM at YouTube scriptwriting, start to finish.
- 1
Real hit videos
We picked real hit YouTube videos from different genres. Each task asks the models to write that video again, and the real script serves as the reference point.
- 2
One tool, many models
Every model wrote inside the TubeLab AI Scriptwriter, the same tool creators use. Same voice training, same research, same steps. Only the model changed.
- 3
Same inputs for everyone
Every model got the same short idea and a voice trained on the original channel, then wrote a full script from scratch. Research was locked per task, so every model worked from the same facts.
- 4
Blind judging
Three AI judges from different companies (OpenAI, Anthropic, Google) grade every script on six writing skills, and models are ranked by those craft scores. They also grade the real video's script the same way, which shows how each script compares with the original. Judges never know which model wrote what, and no judge grades scripts from its own company. The finer scoring rules are in the FAQ.
FAQ
- How is the overall quality score calculated?
- Every script is graded on a six-skill rubric (hook, retention structure, storytelling, voice match, pacing, clarity), and a model's overall score is its average across those six skills and all tasks. Each task is scored by a panel of three judge models from different families (OpenAI, Anthropic, Google); judges never see which model wrote a script, and no judge scores scripts written by its own model family. Tasks a model refuses or fails are scored as the worst craft score any model achieved on that task, so a refusal can never inflate an average. The comparison against the original video is reported separately and does not drive the ranking.
- How do the scripts compare with the real videos?
- Each generated script and the original video's transcript get graded on the same rubric by the same judge. If the generated script scores more than a quarter point higher, it is rated above the original; more than a quarter point lower, below; within the band, it is on par. These comparisons are shown for context and do not affect the ranking, which comes from craft scores alone. Two caveats still apply: this measures what AI judges think of the text, not what audiences would watch, and a transcript strips out the real video's delivery, editing, thumbnail, and proven real-world performance.
- Why did some models not complete every task?
- Claude Fable 5 refused to write two scripts on content-safety grounds, and GPT-5.6 Sol failed to complete the writing flow on one task. These are penalized in the overall score, counted as the worst craft score any model achieved on that task: for a creator, a model that will not write the script has failed it. The craft score covers completed tasks only, and both numbers are shown.
- How are cost and speed measured?
- Cost is the average provider API cost of producing one finished script, including the voice-training and research steps. Speed is the average wall-clock time from brief to finished script.
- Can I write YouTube scripts with these AI models myself?
- Yes. The TubeLab AI Scriptwriter lets you pick the LLM that writes for you, including the top-ranked ones, with a voice trained on your own channel.
- How much do the judges agree with each other?
- On the script-vs-original verdicts, the panel was unanimous 46% of the time, with 52% mean pairwise agreement and a Fleiss' kappa of 0.24 over full three-judge cells (fair agreement beyond chance; many disagreements are win-vs-tie calls near the quarter-point band). Every overall score carries a bootstrap 95% confidence interval, shown on the rankings chart; models with overlapping intervals should be read as ties. The judges have not yet been validated against human ratings or real audience retention.
- Why not just ask the judges to pick the better script?
- We did at first, and the measurements showed it was biased. In forced head-to-head picks the judges never once called a tie, and our most neutral judge (Gemini, with no model in the race) preferred the AI script 76% of the time, more than either judge whose family competes. LLM judges are known to favor LLM-style prose, and the human side competes as a transcript stripped of delivery and visuals. Grading both scripts on the same rubric with a tie band fixed the worst of it: ties reappeared, and the weakest models now lose to the originals, which the old method had hidden.
- Does this pick the best model for scriptwriting in general?
- No. It picks the best model inside the TubeLab Scriptwriter, which is the choice this page is built to help with. Every model ran the same pipeline: the same prompts, the same tools, the same writing steps. A model that fits that pipeline well scores better in it, and a different tool with different prompts could order these models differently. What this measures is each model doing the full job, from research to finished script, not answering one prompt in a vacuum.
- Do different models get scored by different judges?
- Yes. No judge grades scripts from its own company, so Claude scripts are graded by the OpenAI and Google judges, GPT scripts by the Anthropic and Google judges, and models with no judge in the race by all three. We checked whether this skews the board. On the scripts all three judges graded, their average scores sit within about a quarter point of each other on the 10-point scale, and the Anthropic judge turned out to be the most lenient, which means the strictest panel is the one grading Claude models. We also re-ranked every model using only the Google judge, the one judge that scores everyone: the top four positions do not change. A judge panel with no competing families at all is planned for the next update.
- Who runs this benchmark, and what is the conflict of interest?
- TubeLab builds the Scriptwriter used to run every task, and this page promotes it. That conflict cannot be removed, so the design leans on verifiability instead: blind judging by three model families with a self-preference guard, refusals scored as the worst result on the task, published methodology and rubric, and confidence intervals on every score.
- How often is the benchmark updated?
- The benchmark was run in July 2026. We re-run it as significant new models ship; closely ranked models should be read as statistical ties.
Conclusion
What we would pick after judging every script.
Claude Opus 5 is the best AI model for YouTube scriptwriting in this test. It wrote a full script for all 12 tasks, posted the highest craft scores on the board, and judges rated it on par with the original video in most matchups.
The right pick still depends on what you care about. Kimi K3 delivers the most quality per dollar at the lowest price per script. Grok 4.5 balances price, speed, and craft. Claude Fable 5 writes at the same level as Opus 5, but its safety rules refused two topics, and a model that will not write can leave you without a script.
Keep the limits in mind: the judges are AI models, the sample is small, and close scores should be read as ties. Every script on this page came from the TubeLab AI Scriptwriter, so the fastest way to form your own opinion is to pick a model and write with it.
Your subscribers are waiting...
Type an idea. The AI Scriptwriter researches it, plans it, and writes the full script in your channel voice.