This is from Bench's own experience. It is not what any of the frontier models say about themselves. It will change as new models roll out, or as the ones already here get better.
Bench starts with Grok because it's fast. Then the job picks who sits next. One model writes. A different model does the work. A third one tries to break it.
If a word has a star on Glossary / FAQs, it jumps to a plain-English definition.
The short version
- When Bench runs an orchestra, Bench starts with Grok. It's fast. It conducts.
- When Bench puts a new client on their Mac, start with one chat. One problem. Finish it.
- Opus is the model for a normal Tuesday. Fable is the expensive one — use it to build a pipeline (a system that keeps doing the work), not a one-shot (one quick job, then you're done).
- Claude explores and designs. Codex builds, checks, and talks to customers. Grok starts, and it makes video.
- If the chat gets weird or stuck, it's probably out of thinking budget. Start a new one.
Two ways Bench uses the frontier models
Same three companies. Different first move.
When Bench builds
Grok first. Always. It is fast, it stays with you, and it starts the orchestra.
Match the models to the size of the job. A copy change stays in Grok. A real build: Claude writes the plan, Codex does the work, then a second model tries to find the mistakes.
Don't call it done because the first model said it was done.
When Bench starts a new client
One chat. One problem. On their Mac. Watch it finish. Then they earn the second model.
Opus for most days. Fable is rare. A new client does not need an orchestra on day one. Start simple. Like training wheels.
When Bench builds an orchestra
An orchestra is not three chat windows. It is a few models taking turns, each doing the part they're best at. When Bench builds, Grok sits in front.
- Grok starts. Fast. Stays with you. Writes down the job. Calls the next model. Does not decide the final plan.
- Claude plans. Fable when you don't even know what the job is yet. Opus when you already know the shape. Decides what gets built.
- Codex builds. Does the actual work. Tests it. Sticks to the plan.
- A second model checks. An adversarial pass: a second model looks for mistakes.
"I have one chat with Grok. Grok calls Fable. Fable builds. Once Fable's done building it, then I hand it over to Codex. Codex does an adversarial review. That's called an orchestra."
— Bench, during an onboarding session
Match how many models you use to how big the job is. A short copy change stays in Grok. One new feature: Grok starts, Codex builds, Codex checks. A pile of related features: Grok starts, Claude finds the files, Fable or Opus writes the plan, Codex builds each piece. Checkout, legal, or money: the full orchestra. Always.
Don't put Fable on every small task. Don't spend the most thinking on the first try.
The complexity of the task picks the model
Better models give better results. They also take more time and use up more of the monthly thinking budget. You do not put every question on the most expensive model.
"The complexity of the task generally determines which model I use."
— Bench, during an onboarding session
Anthropic (Claude)
| Model | When |
|---|---|
| Fable | Rare. Use it to build a system that keeps doing the work — a video machine for every job, a multi-day plan, a new market. Do not use it to make one video right now. That is an Opus one-shot. |
| Opus | Almost everything. Nearly as good as Fable, and a lot cheaper. When Fable is running low, switch here. This is the model for a normal Tuesday. |
| Sonnet | One quick job. Check an email. Do this one thing. Nothing messy in between. |
| Haiku | A voice model. Not the one you pick for the work. |
"I would say 99.9% of everything you do is going to be Opus. Opus is so good, and it's so much more affordable. There's maybe ten of the things that I do that I would use Fable for."
— Bench, during an onboarding session
OpenAI (ChatGPT / Codex)
| Model | When |
|---|---|
| Sol | The heavy OpenAI model. Hard jobs. Cheaper per run than Fable. Named for the sun — not an agent's soul. |
| Terra | The OpenAI model for most days. More than Luna, less than Sol. |
| Luna | Cheap enough to leave running. A fresh chat with no leftover notes. |
Do not burn the monthly thinking budget to zero. Keep 10–20% in reserve for the job that actually needs it. When Fable is low, switch to Opus.
Why Bench uses all three
You do not need three companies to write an email. You need all three when the job is a finished piece — a photo, a video, and a page.
- Grok starts the orchestra, and makes video
- Claude does layout, type, the page
- ChatGPT makes the still image
Grok starts the work. It does not decide the final plan. Claude decides the plan. Codex builds it.
Claude explores. Codex runs.
This is the pair people get backwards. They build the new thing in Codex, then wonder why it cannot see the whole job.
"Codex is a great builder, but it is not a visionary. The experimental new stuff that you're building out — Claude is the stuff that you want to build with."
— Bench, during an onboarding session
Never put Claude in front of a customer. Never ask Codex to figure out what to build. Never let Grok decide the final plan.
Thinking is a budget
Which model you pick, and how hard you tell it to think, are the same kind of choice. More thinking budget means harder problems, slower answers, and more of the monthly allowance used up.
High is enough for most work. Ultra is for a huge project. If the agent gets weird or stuck halfway through a big job, this chat is out of thinking budget. Start a new one.
Claude is a designer. It is not a camera.
Someone says "make me a door hanger." Claude tries to invent the picture out of shapes. The page looks fake.
Start with the problem, not the artifact. Get the photos from the jobs — AccuLynx, the drive, the download — or generate them in ChatGPT or Gemini. Hand those to Claude.
Sit down and use it
The ladder is a starting map. It is not a religion. Start with the problem in front of you.
Study the words on Glossary / FAQs.
Subscribe
Bench dispatches new field intelligence, operator notes, and agent-authored research here.
