What If Claude Fable 5 Goes Away? Three Controllers Compared
With three controllers in play, we line up all nine deliverables and compare.
Previously: Can Claude Opus Run the Lane? Opus vs Fable as Controller
This is the final installment of the “Claude vs Codex Arena” series. So far we have compared Claude Code (model: Fable 5) and Codex (model: GPT-5.6-Sol, reasoning effort: ultra) as controllers, and last time we measured the gap against a third controller candidate, Opus 4.8. In this final piece I walk through how the three-way contest played out, in order, and then hand the lessons from the whole experiment over to our collaborative development engine, ModelOrcs.
This article records one task with a single run per setup (n=1). It is not evidence about how these models compare in general.
How we got to a three-way contest (recap)
The details are in the previous installment, so here are just the essentials.
The original motivation was the possibility that Fable 5 would drop out of the flat-rate subscription plans — at the time of the experiment, the situation read as “it may become unavailable after July 19.” Scouting a fallback, I decided to try Opus 4.8 as controller. Along the way came the swapped-reviewer incident, where a fresh Fable session brought in for review switched to Opus partway through a turn. According to Anthropic’s official support article, Fable has safeguards, and for certain content the model is designed to switch to Opus automatically, even mid-turn (as documented when we checked in July 2026). Our fail-closed verifier detected the mixed models and invalidated the review. Both the safeguards and the verifier worked as designed — and the episode doubled, unintentionally, as a preview of Opus standing in for Fable.
The comparison ran in two stages. In a replay that changed only the merge step between two finished candidates, the Opus merge scored 63.63 against the Fable merge’s 47.84 (+15.79). But that compares merge duty alone and leaves most of a controller’s job unmeasured. So in Round 4 we handed the entire lane to Opus, starting from the public baseline. After a fix for a sandbox that would not come up, and two capacity interruptions with auto-resume, the lane ran to completion.
That gave us all nine deliverables: 3 controllers × 3 outputs. Here is how the three lanes progressed, on one chart.
One caution about reading the time. Most of the Opus lane’s 14 hours and 21 minutes was failures, approvals, and waiting on supply; the processes were actively working for a combined 2 hours and 4 minutes. You cannot read the difference in wall time as a difference in thinking speed. What you can read from it is whether a lane recovers safely from failure and finishes — a legitimate evaluation criterion when you are handing over a long piece of work.
All nine side by side — which combination looks best
Once all nine deliverables (3 lanes × 3 outputs) were in, we froze them and rescored every candidate on identical fresh data generated from a new seed (evaluation run eval4-20260718-r3). The ranking:
| Rank | Controller lane | Deliverable | Score |
|---|---|---|---|
| 1 | Fable | Codex solo | 65.07 |
| 2 | Codex | Codex Mix | 58.32 |
| 3 | Codex | Codex solo | 57.07 |
| 4 | Opus | Codex solo | 55.89 |
| 5 | Opus | Opus Mix | 48.15 |
| 6 | Fable | Fable Mix | 47.84 |
| 7 | Codex | Claude/Fable solo | 45.36 |
| 8 | Fable | Claude/Fable solo | 44.30 |
| 9 | Opus | Claude/Fable solo | 42.92 |
This table is the answer key for the whole series. Reading across all the combinations, here is what stands out.
- The most promising setup is a Fable controller with a Codex worker (65.07). The top of the field came from that pairing.
- The Codex-controlled lane is the runner-up. Codex Mix at 58.32 and Codex solo at 57.07 took second and third overall.
- The Opus-controlled lane sits in the middle. Its best score again came from a Codex worker (55.89). It did not reach what the Fable-led lane produced, but it did show the operational strength to carry a lane to completion.
- And under every controller, the Codex worker was the backbone of the result. The Claude/Fable worker’s solo output landed in the bottom group in all three lanes.
What we wanted to see was never proof of whether Opus is better — it was what changes when you replace the controller. The answer: the ranking did not reorder itself, and what drove the outcome was neither the controller alone nor the worker alone, but the combination.
This is not a general statement that Opus is worse than Fable. It is one task, one set of evaluation data, and a single run per lane, with some instruction conditions still not perfectly aligned (disclosed where relevant). On top of that, every lane kept its workers fixed to Fable and Codex, so a configuration like “Opus controller with Opus workers” was never tested. The claim stops here: this time, how the controller and workers were paired mattered more than which controller was in the seat.
One more thing about the numbers. Precision across the nine deliverables ranged from 5.75% to 11.32% (measured against a synthetic oracle — machine-generated ground truth, not production accuracy). Even the top-ranked deliverable is at 11.32%. Producing a ranking is a separate matter from being able to hand off sending email, cutting prices, or placing orders unattended. Every one of these deliverables is best understood as a candidate-shortlisting assistant with a human reviewer in the loop.
What we learned across the whole experiment
Narrowed down, the series taught us five things.
- AI output needs provenance checks. The swapped-reviewer incident was only caught because a fail-closed verifier was in place. Record the model, the inputs, and every stop and restart, and halt on doubt when an assumption breaks. Halting is not a failure; it is the feature that keeps you from proceeding with your comparison conditions already broken.
- A convenient shortcut experiment can quietly swap out the question. The +15.79 was a correct number, but it answered a different question than the one we started with. Generating a standalone implementation, fusing two candidates, and carrying a full lane through failures are three different abilities — and you must not put a label on a measurement that is bigger than what you measured.
- Keep time and score in separate ledgers. Wall time mixes in failures, waiting, and approvals. Without separating them, you will mistake a lucky lane for a fast model.
- A merge is decided by which candidate you build on, and the merged result is a third deliverable that has to be re-tested. Adding the strengths of two candidates does not automatically produce something better, and the fact that both originals passed their tests does not let you skip testing the merge.
- Design interruptions as normal operation, not as exceptions. Only when you have resumable state, defined conditions for retrying, for stopping, and for handing back to a human can you trust an AI with a long piece of work. And every configuration in this experiment produced quality that assumes human review.
What to do from July 20: three options for staffing the controller
While this experiment was running, Fable 5 looked likely to drop out of the flat-rate subscription plans, which would have made the setup that starred as controller in this series unusable as-is. In the end the treatment differs by plan (see below), but the underlying question — what do you do when a specific model becomes unavailable to you — still stands. Here are the options, within what this experiment can support.
Update: the outcome was unknown while these experiments were running. According to Anthropic's official help center, from July 20, 2026 Fable 5 is included on Max plans and premium seats on Team plans, where up to 50% of the weekly usage limit can go to Fable 5 at no extra cost (this is not additive — other models draw from the same limit). On Pro plans and standard Team seats, Fable 5 runs on pay-as-you-go usage credits instead. Fable 5 also consumes those limits faster than other Claude models. So the premise that Fable would become entirely unavailable did not hold, depending on the plan. The experiments here were nonetheless run under the uncertainty that existed beforehand. Terms can change, so always confirm current conditions against official sources (as of July 20, 2026).
A further update: on July 24, 2026, Anthropic announced Claude Opus 5, the successor to Opus 4.8. According to the official announcement, it approaches Fable 5's frontier intelligence at roughly half the price. Every experiment and score in this series used Opus 4.8. Always check the current model lineup and terms against official sources (as of July 27, 2026).
- Keep using Fable. On this one task, the Fable-controlled lane produced the best deliverable. On a Max plan that means working within the weekly limit; on Pro it means pay-as-you-go credits. Either way it is the quality-first option, with consumption and cost measured as you go.
- Move the controller to Opus or Codex. Opus demonstrated safe stop-and-resume through failures and finished its lane; Codex produced the highest-scoring Mix. But these results come from a single task, and switching is no guarantee of the same quality.
- Make the controller swappable. Stop depending on any specific model, and build a setup where both the controller and the workers can be replaced. Just before this article went out, a new candidate appeared in the form of Opus 5 — the lineup of models will keep changing. That is the path we took, and it leads to ModelOrcs.
And so, ModelOrcs v0.1
This series was always a run-up to building “an engine that runs Claude and Codex in parallel, then adopts or merges the better result based on machine scoring and adjudication.” ModelOrcs v0.1, which reflects the lessons above, now exists. The repository is currently private; if we finish preparing it for release — auditing and cleaning up the code and history — we are considering publishing it as open source.
The design draws its boundaries exactly where this experiment hurt. Deterministic control — worktree isolation, parallel execution, timeouts, retries, state persistence, score aggregation — is handled by the program, and the LLM is left with generation, critique, and merging only. A single cycle looks like this.
- Split off separate worktrees for Claude and for Codex from the same base SHA, and generate candidates in parallel
- Run the shared tests and machine scoring first, then pass the results to a structured adjudication with candidate names withheld
- Have the engine mechanically execute whatever the adjudication returns — adopt candidate A, adopt candidate B, merge, or stop for a human — and always log the reasoning behind the decision
- Re-run the full test suite against the adopted or merged result, and hand back to a human on any concern. Merging to the main branch is done by a human
Note that withholding candidate names is not a guarantee of fairness; it is one measure for reducing bias in adjudication (LLM judgments can still carry biases such as position effects). We are not detailing the scoring formula or adjudication criteria here.
v0.1 was just born, and there is nothing we can say about its performance yet. Nor did the results of this series prove anything about the performance of a collaborative approach. The experiments to establish that will be run separately, with the same kind of evidence trail this series kept.
Closing the series
From the first installment, where the same question produced completely different proposals, to this finale, where the judges themselves changed, one thing stayed constant: build the means of verification before you trust what an AI produces. Adding more AI does not by itself produce collaboration. Who gets which job, what yardstick you compare them on, where you stop, and what the human takes on. Every one of these models is capable, and not one of them is grounds for trust on its own. Structure and verification are what determine quality — n=1, but that is what we at MIF came away believing.
The era of choosing between AIs is giving way to an era of designing how you use them. We hope this series helps with that. Thank you for reading.
This series (6 parts)
- Same Question, Different Proposals
- 100 Connect Four Games Head-to-Head
- Claude Code vs Codex on Real EC Data
- Does Merging Beat the Best Solo Run?
- Can Claude Opus Run the Lane? Opus vs Fable as Controller
- Three Controllers Compared, and ModelOrcs (this article)