Rankings
The top 5 AI models for YouTube scriptwriting, by craft score. Judges graded every script on six writing skills; how each script fared against the real video it was based on is shown separately. The thin lines show the range each score could fall in. Opus 5 leads, but the group behind it is close.
LLM Leaderboard
Every LLM we benchmarked for scriptwriting: craft score, relative cost per script, and writing time. Rank the board by a single skill to see which model writes the best hooks, storytelling, or voice match, and open a row for the full six-skill breakdown.
| Model | Overall | Cost | Speed | ||
|---|---|---|---|---|---|
| 1 | Claude Opus 5Anthropic | 87.9% | 1.8× | ~8.3 min | |
Best overall. It wrote a full script for all 12 tasks, earned the highest craft scores in the test, and was the most consistent model from task to task.
| |||||
| 2 | Claude Fable 5.1Anthropic | 85.7% | 2.7× | ~6.7 min | |
| 3 | Claude Opus 4.8Anthropic | 84.4% | 1.4× | ~6.4 min | |
| 4 | K3Kimi K3Moonshot AI | 83.9% | 1× | ~8.6 min | |
| 5 | GLMGLM 5.3Zhipu AI | 83.6% | 0.4× | ~9.2 min | |
| 6 | XGrok 4.5xAI | 82.8% | 1.3× | ~6.8 min | |
| 7 | Claude Fable 5Anthropic | 82.2% | 3× | ~7.5 min | |
| 8 | XGrok 4.6xAI | 80.4% | 1.4× | ~9.8 min | |
| 9 | GPT-5.5OpenAI | 80.3% | 1.1× | ~4.7 min | |
| 10 | Claude Sonnet 5Anthropic | 79.3% | 0.7× | ~7.5 min | |
| 11 | MMuse Spark 1.1Meta | 79% | 1.1× | ~4.5 min | |
| 12 | GGemini 3.8 FlashGoogle | 78.5% | 0.17× | ~1.8 min | |
| 13 | GLMGLM 5.2Zhipu AI | 77.8% | 0.6× | ~14.6 min | |
| 14 | DSDeepSeek V4 Pro (0813)DeepSeek | 76.5% | 0.4× | ~5.3 min | |
| 15 | DSDeepSeek V4 ProDeepSeek | 76.4% | 0.05× | ~10.6 min | |
| 16 | DSDeepSeek V4 FlashDeepSeek | 75.9% | 0.05× | ~8.9 min | |
| 17 | Claude Sonnet 4.6Anthropic | 75.4% | 0.7× | ~7.3 min | |
| 18 | GPT-5.6 SolOpenAI | 75.3% | 1× | ~4.1 min | |
| 19 | GGemini 3.6 FlashGoogle | 75% | 0.4× | ~3.9 min | |
| 20 | MMuse Spark 1.3Meta | 72.8% | 0.3× | ~4.3 min | |
| 21 | K2Kimi K2.5Moonshot AI | 71.3% | 0.13× | ~6.4 min | |
| 22 | GPT-6 AstraOpenAI | 56.7% | 1.8× | ~2.8 min | |
Cost vs performance
Each AI model's craft score against the relative cost of one script. Higher up means better. Further left means cheaper.
Key takeaways
Which AI model to pick for YouTube scriptwriting, based on what you care about most: quality, reliability, cost, or speed.
Best overall: Claude Opus 5
A clear number one. It wrote a full script for all 12 tasks, got the highest craft scores from the judges, and was the most consistent model from task to task.
Best value: Kimi K3
Fourth overall at about half the cost of Opus 5. GLM 5.3 scores nearly as well for less than half of that, but it stalled on one task and never delivered.
Best writing, with a catch: Claude Fable 5
Craft scores tied with Opus 5, but its strict safety rules made it refuse two of the 12 scripts: a disturbing sea creatures topic and a CIA hacking story. A refused task gets the worst score of that round, which drops it to 7th, and its successor Fable 5.1 wrote every script.
Cheapest: DeepSeek V4 Pro
A full script for about a twentieth of what the mid-priced models cost, with middle-of-the-board craft. It is slow, at over 10 minutes per script, and its Flash sibling matches the price but stalled on one task.
Fastest: Gemini 3.8 Flash
Under two minutes per script, among the cheapest on the board, and it finished all 12 tasks. Judges still preferred the original video in 8 of them, so it lands mid-board on craft.
Surprising results
The findings we did not see coming.
GLM 5.3 went from 13th to 5th
GLM 5.2 sits 13th on this board. The point release right after it lands 5th, level with Kimi K3 and Claude Opus 4.8, and it costs less than half of anything else in the top five. One small version bump moved it past eight other models, which is a good reason to re-test a model you ruled out a few months ago.
OpenAI's models: great at code, weak at scripts
The two newest OpenAI models landed near the bottom. GPT-5.6 Sol, a top model on coding benchmarks, sits 18th with a failed task and a weak voice match. GPT-6 Astra, the newest flagship, came 22nd of 22, more than fourteen points behind the next model and the only entry the judges rated below the original video on every task: the scripts were on length and on time, but read like briefings, with hypothetical examples instead of stories and little of the channel's voice. The older GPT-5.5 is the best of the three, at 9th. Models tuned that hard for reasoning and code seem to pay for it in creative work. One caveat: the Scriptwriter runs OpenAI models at their lowest reasoning setting to keep costs down, and a higher setting might write differently, at a higher price.
Grok 4.5 is the dark horse
Nobody talks about Grok for YouTube scripts, yet it landed 6th, ahead of GPT-5.5 and three of the six Claude models. It matched channel voices better than most, at a middle price, and the newer Grok 4.6 landed two places below it.
Fable 5: Refused writing scripts
Fable 5 scored level with the winner and still fell to 7th, only because it refused two topics: a video about disturbing deep-sea creatures and one about the teens who talked their way into the CIA. A model you cannot count on loses points, no matter how well it writes. Fable 5.1 fixed that: it wrote every script and sits 2nd.
DeepSeek V4 costs pennies on the dollar
Both DeepSeek V4 models write a full script for eight to fifty-five times less than the top five, and still finish mid-board, ahead of Claude Sonnet 4.6. The catch is speed for Pro and one stalled task for Flash.
Kimi K2.5 is not a smaller K3
The two Moonshot models sit at opposite ends of the board: K3 is 4th, K2.5 is 21st of 22, more than twelve points behind it. The family name tells you nothing about the writing.
Your next video starts here.
Write yours with the same tool. Give the AI Scriptwriter an idea and get back a full script, researched and written in your channel voice.
Craft comparison
How the top 5 AI scriptwriting models score on the six craft skills: hook, retention, storytelling, voice match, pacing, and clarity. Judges scored each from 1 to 10. The other models are in the leaderboard rows.
Generated scripts
Each task is based on a real hit video. Open one to read the full script each LLM wrote and see its scores. All scripts come from TubeLab's AI Scriptwriter.
How we tested it
How we benchmarked each LLM at YouTube scriptwriting, start to finish.
- 1
Real hit videos
We picked real hit YouTube videos from different genres. Each task asks the models to write that video again, and the real script serves as the reference point.
- 2
One tool, many models
Every model wrote inside the TubeLab AI Scriptwriter, the same tool creators use. Same voice training, same research, same steps. Only the model changed.
- 3
Same inputs for everyone
Every model got the same short idea and a voice trained on the original channel, then wrote a full script from scratch. Research was locked per task, so every model worked from the same facts.
- 4
Blind judging
Three AI judges from different companies (OpenAI, Anthropic, Google) grade every script on six writing skills, and models are ranked by those craft scores. They also grade the real video's script the same way, which shows how each script compares with the original. Judges never know which model wrote what, and no judge grades scripts from its own company. The finer scoring rules are in the FAQ.
FAQ
- How is the overall quality score calculated?
- Every script is graded on a six-skill rubric (hook, retention structure, storytelling, voice match, pacing, clarity), and a model's overall score is its average across those six skills and all tasks. Each task is scored by a panel of three judge models from different families (OpenAI, Anthropic, Google); judges never see which model wrote a script, and no judge scores scripts written by its own model family. Tasks a model refuses or fails are scored as the worst craft score any model achieved on that task, so a refusal can never inflate an average. The comparison against the original video is reported separately and does not drive the ranking.
- How do the scripts compare with the real videos?
- Each generated script and the original video's transcript get graded on the same rubric by the same judge. If the generated script scores more than a quarter point higher, it is rated above the original; more than a quarter point lower, below; within the band, it is on par. These comparisons are shown for context and do not affect the ranking, which comes from craft scores alone. Two caveats still apply: this measures what AI judges think of the text, not what audiences would watch, and a transcript strips out the real video's delivery, editing, thumbnail, and proven real-world performance.
- Why did some models not complete every task?
- Claude Fable 5 refused to write two scripts on content-safety grounds, GPT-5.6 Sol failed to complete the writing flow on one task, DeepSeek V4 Flash and GLM 5.3 each stalled mid-write on one task and never delivered a script, and Grok 4.6 and DeepSeek V4 Pro (0813) ran out of steps mid-write on two tasks and one task respectively. These are penalized in the overall score, counted as the worst craft score any model achieved on that task: for a creator, a model that will not write the script has failed it. The craft score covers completed tasks only, and both numbers are shown.
- How are cost and speed measured?
- Cost is the average provider API cost of producing one finished script, including the voice-training and research steps. We show it as a multiple of a fixed baseline rather than in dollars: provider prices change too often for exact figures to stay accurate, and the ratio between models is what matters for the choice. Speed is the average wall-clock time from brief to finished script.
- Can I write YouTube scripts with these AI models myself?
- Yes. The TubeLab AI Scriptwriter lets you pick the LLM that writes for you, including the top-ranked ones, with a voice trained on your own channel.
- How much do the judges agree with each other?
- On the script-vs-original verdicts, the panel was unanimous 45% of the time, with 54% mean pairwise agreement and a Fleiss' kappa of 0.27 over full three-judge cells (fair agreement beyond chance; many disagreements are win-vs-tie calls near the quarter-point band). Every overall score carries a bootstrap 95% confidence interval, shown on the rankings chart; models with overlapping intervals should be read as ties. The judges have not yet been validated against human ratings or real audience retention.
- Why not just ask the judges to pick the better script?
- We did at first, and the measurements showed it was biased. In forced head-to-head picks the judges never once called a tie, and our most neutral judge (Gemini, with no model in the race) preferred the AI script 76% of the time, more than either judge whose family competes. LLM judges are known to favor LLM-style prose, and the human side competes as a transcript stripped of delivery and visuals. Grading both scripts on the same rubric with a tie band fixed the worst of it: ties reappeared, and the weakest models now lose to the originals, which the old method had hidden.
- Does this pick the best model for scriptwriting in general?
- No. It picks the best model inside the TubeLab Scriptwriter, which is the choice this page is built to help with. Every model ran the same pipeline: the same prompts, the same tools, the same writing steps. A model that fits that pipeline well scores better in it, and a different tool with different prompts could order these models differently. What this measures is each model doing the full job, from research to finished script, not answering one prompt in a vacuum.
- Do different models get scored by different judges?
- Yes. No judge grades scripts from its own company, so Claude scripts are graded by the OpenAI and Google judges, GPT scripts by the Anthropic and Google judges, and models with no judge in the race by all three. We checked whether this skews the board. On the scripts all three judges graded, their averages sit within about half a point of each other on the 10-point scale, and the Anthropic judge is the most lenient, which means the strictest panel is the one grading Claude models. We also re-ranked every model using only the Google judge, which scores every model except Gemini 3.6 Flash: the top five positions do not change. A judge panel with no competing families at all is planned for the next update.
- Who runs this benchmark, and what is the conflict of interest?
- TubeLab builds the Scriptwriter used to run every task, and this page promotes it. That conflict cannot be removed, so the design leans on verifiability instead: blind judging by three model families with a self-preference guard, refusals scored as the worst result on the task, published methodology and rubric, and confidence intervals on every score.
- How often is the benchmark updated?
- The benchmark was run in July 2026, updated in August 2026 with GLM 5.3, Grok 4.6 and DeepSeek V4 Pro (0813), and again in September 2026 with Claude Fable 5.1, Gemini 3.8 Flash, Muse Spark 1.3 and GPT-6 Astra. We re-run it as significant new models ship; closely ranked models should be read as statistical ties.
Conclusion
What we would pick after judging every script.
Claude Opus 5 is the best AI model for YouTube scriptwriting in this test. It wrote a full script for all 12 tasks, posted the highest craft scores on the board, and judges rated it on par with the original video in most matchups.
The right pick still depends on what you care about. Kimi K3 delivers the most quality per dollar of the models that finished every task, at about half the price of Opus 5. DeepSeek V4 Pro writes a script for a twentieth of the usual cost and lands mid-board, if you can wait ten minutes for it. Grok 4.5 balances price, speed, and craft. Claude Fable 5.1 writes at the same level as Opus 5 and, unlike Fable 5, wrote every script without a refusal, but it costs about half again as much.
Keep the limits in mind: the judges are AI models, the sample is small, and close scores should be read as ties. Every script on this page came from the TubeLab Scriptwriter, so the fastest way to form your own opinion is to pick a model and write with it.
Your subscribers are waiting...
Type an idea. The Scriptwriter researches it, plans it, and writes the full script in your channel voice.
Rankings
The top 5 AI models for YouTube scriptwriting, by craft score. Judges graded every script on six writing skills; how each script fared against the real video it was based on is shown separately. The thin lines show the range each score could fall in. Opus 5 leads, but the group behind it is close.
LLM Leaderboard
Every LLM we benchmarked for scriptwriting: craft score, relative cost per script, and writing time. Rank the board by a single skill to see which model writes the best hooks, storytelling, or voice match, and open a row for the full six-skill breakdown.
| Model | Overall | Cost | Speed | ||
|---|---|---|---|---|---|
| 1 | Claude Opus 5Anthropic | 87.9% | 1.8× | ~8.3 min | |
Best overall. It wrote a full script for all 12 tasks, earned the highest craft scores in the test, and was the most consistent model from task to task.
| |||||
| 2 | Claude Fable 5.1Anthropic | 85.7% | 2.7× | ~6.7 min | |
| 3 | Claude Opus 4.8Anthropic | 84.4% | 1.4× | ~6.4 min | |
| 4 | K3Kimi K3Moonshot AI | 83.9% | 1× | ~8.6 min | |
| 5 | GLMGLM 5.3Zhipu AI | 83.6% | 0.4× | ~9.2 min | |
| 6 | XGrok 4.5xAI | 82.8% | 1.3× | ~6.8 min | |
| 7 | Claude Fable 5Anthropic | 82.2% | 3× | ~7.5 min | |
| 8 | XGrok 4.6xAI | 80.4% | 1.4× | ~9.8 min | |
| 9 | GPT-5.5OpenAI | 80.3% | 1.1× | ~4.7 min | |
| 10 | Claude Sonnet 5Anthropic | 79.3% | 0.7× | ~7.5 min | |
| 11 | MMuse Spark 1.1Meta | 79% | 1.1× | ~4.5 min | |
| 12 | GGemini 3.8 FlashGoogle | 78.5% | 0.17× | ~1.8 min | |
| 13 | GLMGLM 5.2Zhipu AI | 77.8% | 0.6× | ~14.6 min | |
| 14 | DSDeepSeek V4 Pro (0813)DeepSeek | 76.5% | 0.4× | ~5.3 min | |
| 15 | DSDeepSeek V4 ProDeepSeek | 76.4% | 0.05× | ~10.6 min | |
| 16 | DSDeepSeek V4 FlashDeepSeek | 75.9% | 0.05× | ~8.9 min | |
| 17 | Claude Sonnet 4.6Anthropic | 75.4% | 0.7× | ~7.3 min | |
| 18 | GPT-5.6 SolOpenAI | 75.3% | 1× | ~4.1 min | |
| 19 | GGemini 3.6 FlashGoogle | 75% | 0.4× | ~3.9 min | |
| 20 | MMuse Spark 1.3Meta | 72.8% | 0.3× | ~4.3 min | |
| 21 | K2Kimi K2.5Moonshot AI | 71.3% | 0.13× | ~6.4 min | |
| 22 | GPT-6 AstraOpenAI | 56.7% | 1.8× | ~2.8 min | |
Cost vs performance
Each AI model's craft score against the relative cost of one script. Higher up means better. Further left means cheaper.
Key takeaways
Which AI model to pick for YouTube scriptwriting, based on what you care about most: quality, reliability, cost, or speed.
Best overall: Claude Opus 5
A clear number one. It wrote a full script for all 12 tasks, got the highest craft scores from the judges, and was the most consistent model from task to task.
Best value: Kimi K3
Fourth overall at about half the cost of Opus 5. GLM 5.3 scores nearly as well for less than half of that, but it stalled on one task and never delivered.
Best writing, with a catch: Claude Fable 5
Craft scores tied with Opus 5, but its strict safety rules made it refuse two of the 12 scripts: a disturbing sea creatures topic and a CIA hacking story. A refused task gets the worst score of that round, which drops it to 7th, and its successor Fable 5.1 wrote every script.
Cheapest: DeepSeek V4 Pro
A full script for about a twentieth of what the mid-priced models cost, with middle-of-the-board craft. It is slow, at over 10 minutes per script, and its Flash sibling matches the price but stalled on one task.
Fastest: Gemini 3.8 Flash
Under two minutes per script, among the cheapest on the board, and it finished all 12 tasks. Judges still preferred the original video in 8 of them, so it lands mid-board on craft.
Surprising results
The findings we did not see coming.
GLM 5.3 went from 13th to 5th
GLM 5.2 sits 13th on this board. The point release right after it lands 5th, level with Kimi K3 and Claude Opus 4.8, and it costs less than half of anything else in the top five. One small version bump moved it past eight other models, which is a good reason to re-test a model you ruled out a few months ago.
OpenAI's models: great at code, weak at scripts
The two newest OpenAI models landed near the bottom. GPT-5.6 Sol, a top model on coding benchmarks, sits 18th with a failed task and a weak voice match. GPT-6 Astra, the newest flagship, came 22nd of 22, more than fourteen points behind the next model and the only entry the judges rated below the original video on every task: the scripts were on length and on time, but read like briefings, with hypothetical examples instead of stories and little of the channel's voice. The older GPT-5.5 is the best of the three, at 9th. Models tuned that hard for reasoning and code seem to pay for it in creative work. One caveat: the Scriptwriter runs OpenAI models at their lowest reasoning setting to keep costs down, and a higher setting might write differently, at a higher price.
Grok 4.5 is the dark horse
Nobody talks about Grok for YouTube scripts, yet it landed 6th, ahead of GPT-5.5 and three of the six Claude models. It matched channel voices better than most, at a middle price, and the newer Grok 4.6 landed two places below it.
Fable 5: Refused writing scripts
Fable 5 scored level with the winner and still fell to 7th, only because it refused two topics: a video about disturbing deep-sea creatures and one about the teens who talked their way into the CIA. A model you cannot count on loses points, no matter how well it writes. Fable 5.1 fixed that: it wrote every script and sits 2nd.
DeepSeek V4 costs pennies on the dollar
Both DeepSeek V4 models write a full script for eight to fifty-five times less than the top five, and still finish mid-board, ahead of Claude Sonnet 4.6. The catch is speed for Pro and one stalled task for Flash.
Kimi K2.5 is not a smaller K3
The two Moonshot models sit at opposite ends of the board: K3 is 4th, K2.5 is 21st of 22, more than twelve points behind it. The family name tells you nothing about the writing.
Your next video starts here.
Write yours with the same tool. Give the AI Scriptwriter an idea and get back a full script, researched and written in your channel voice.
Craft comparison
How the top 5 AI scriptwriting models score on the six craft skills: hook, retention, storytelling, voice match, pacing, and clarity. Judges scored each from 1 to 10. The other models are in the leaderboard rows.
Generated scripts
Each task is based on a real hit video. Open one to read the full script each LLM wrote and see its scores. All scripts come from TubeLab's AI Scriptwriter.
How we tested it
How we benchmarked each LLM at YouTube scriptwriting, start to finish.
- 1
Real hit videos
We picked real hit YouTube videos from different genres. Each task asks the models to write that video again, and the real script serves as the reference point.
- 2
One tool, many models
Every model wrote inside the TubeLab AI Scriptwriter, the same tool creators use. Same voice training, same research, same steps. Only the model changed.
- 3
Same inputs for everyone
Every model got the same short idea and a voice trained on the original channel, then wrote a full script from scratch. Research was locked per task, so every model worked from the same facts.
- 4
Blind judging
Three AI judges from different companies (OpenAI, Anthropic, Google) grade every script on six writing skills, and models are ranked by those craft scores. They also grade the real video's script the same way, which shows how each script compares with the original. Judges never know which model wrote what, and no judge grades scripts from its own company. The finer scoring rules are in the FAQ.
FAQ
- How is the overall quality score calculated?
- Every script is graded on a six-skill rubric (hook, retention structure, storytelling, voice match, pacing, clarity), and a model's overall score is its average across those six skills and all tasks. Each task is scored by a panel of three judge models from different families (OpenAI, Anthropic, Google); judges never see which model wrote a script, and no judge scores scripts written by its own model family. Tasks a model refuses or fails are scored as the worst craft score any model achieved on that task, so a refusal can never inflate an average. The comparison against the original video is reported separately and does not drive the ranking.
- How do the scripts compare with the real videos?
- Each generated script and the original video's transcript get graded on the same rubric by the same judge. If the generated script scores more than a quarter point higher, it is rated above the original; more than a quarter point lower, below; within the band, it is on par. These comparisons are shown for context and do not affect the ranking, which comes from craft scores alone. Two caveats still apply: this measures what AI judges think of the text, not what audiences would watch, and a transcript strips out the real video's delivery, editing, thumbnail, and proven real-world performance.
- Why did some models not complete every task?
- Claude Fable 5 refused to write two scripts on content-safety grounds, GPT-5.6 Sol failed to complete the writing flow on one task, DeepSeek V4 Flash and GLM 5.3 each stalled mid-write on one task and never delivered a script, and Grok 4.6 and DeepSeek V4 Pro (0813) ran out of steps mid-write on two tasks and one task respectively. These are penalized in the overall score, counted as the worst craft score any model achieved on that task: for a creator, a model that will not write the script has failed it. The craft score covers completed tasks only, and both numbers are shown.
- How are cost and speed measured?
- Cost is the average provider API cost of producing one finished script, including the voice-training and research steps. We show it as a multiple of a fixed baseline rather than in dollars: provider prices change too often for exact figures to stay accurate, and the ratio between models is what matters for the choice. Speed is the average wall-clock time from brief to finished script.
- Can I write YouTube scripts with these AI models myself?
- Yes. The TubeLab AI Scriptwriter lets you pick the LLM that writes for you, including the top-ranked ones, with a voice trained on your own channel.
- How much do the judges agree with each other?
- On the script-vs-original verdicts, the panel was unanimous 45% of the time, with 54% mean pairwise agreement and a Fleiss' kappa of 0.27 over full three-judge cells (fair agreement beyond chance; many disagreements are win-vs-tie calls near the quarter-point band). Every overall score carries a bootstrap 95% confidence interval, shown on the rankings chart; models with overlapping intervals should be read as ties. The judges have not yet been validated against human ratings or real audience retention.
- Why not just ask the judges to pick the better script?
- We did at first, and the measurements showed it was biased. In forced head-to-head picks the judges never once called a tie, and our most neutral judge (Gemini, with no model in the race) preferred the AI script 76% of the time, more than either judge whose family competes. LLM judges are known to favor LLM-style prose, and the human side competes as a transcript stripped of delivery and visuals. Grading both scripts on the same rubric with a tie band fixed the worst of it: ties reappeared, and the weakest models now lose to the originals, which the old method had hidden.
- Does this pick the best model for scriptwriting in general?
- No. It picks the best model inside the TubeLab Scriptwriter, which is the choice this page is built to help with. Every model ran the same pipeline: the same prompts, the same tools, the same writing steps. A model that fits that pipeline well scores better in it, and a different tool with different prompts could order these models differently. What this measures is each model doing the full job, from research to finished script, not answering one prompt in a vacuum.
- Do different models get scored by different judges?
- Yes. No judge grades scripts from its own company, so Claude scripts are graded by the OpenAI and Google judges, GPT scripts by the Anthropic and Google judges, and models with no judge in the race by all three. We checked whether this skews the board. On the scripts all three judges graded, their averages sit within about half a point of each other on the 10-point scale, and the Anthropic judge is the most lenient, which means the strictest panel is the one grading Claude models. We also re-ranked every model using only the Google judge, which scores every model except Gemini 3.6 Flash: the top five positions do not change. A judge panel with no competing families at all is planned for the next update.
- Who runs this benchmark, and what is the conflict of interest?
- TubeLab builds the Scriptwriter used to run every task, and this page promotes it. That conflict cannot be removed, so the design leans on verifiability instead: blind judging by three model families with a self-preference guard, refusals scored as the worst result on the task, published methodology and rubric, and confidence intervals on every score.
- How often is the benchmark updated?
- The benchmark was run in July 2026, updated in August 2026 with GLM 5.3, Grok 4.6 and DeepSeek V4 Pro (0813), and again in September 2026 with Claude Fable 5.1, Gemini 3.8 Flash, Muse Spark 1.3 and GPT-6 Astra. We re-run it as significant new models ship; closely ranked models should be read as statistical ties.
Conclusion
What we would pick after judging every script.
Claude Opus 5 is the best AI model for YouTube scriptwriting in this test. It wrote a full script for all 12 tasks, posted the highest craft scores on the board, and judges rated it on par with the original video in most matchups.
The right pick still depends on what you care about. Kimi K3 delivers the most quality per dollar of the models that finished every task, at about half the price of Opus 5. DeepSeek V4 Pro writes a script for a twentieth of the usual cost and lands mid-board, if you can wait ten minutes for it. Grok 4.5 balances price, speed, and craft. Claude Fable 5.1 writes at the same level as Opus 5 and, unlike Fable 5, wrote every script without a refusal, but it costs about half again as much.
Keep the limits in mind: the judges are AI models, the sample is small, and close scores should be read as ties. Every script on this page came from the TubeLab Scriptwriter, so the fastest way to form your own opinion is to pick a model and write with it.
Your subscribers are waiting...
Type an idea. The Scriptwriter researches it, plans it, and writes the full script in your channel voice.