Home » Blog » Word on the Future » Claude or Codex? WordPress doesn’t care

Claude or Codex? WordPress doesn’t care

I often get asked, "Claude or Codex?"

For larger companies, this decision matters. You’ll go through procurement, approvals, and more. Yet it probably matters less than people think. Each frontier model has its edge, but that edge comes and goes.

After the billions of tokens I’ve run through, here’s my honest take: you’re fine with whatever you have.

Here’s why…

In 2009, a team published results from a surgical safety study across eight hospitals and 7,688 patients. Complications fell from 11% to 7%, deaths from 1.5% to 0.8%, with only one change. The surgeons who left on Friday were the same ones who came back on Monday. A simple checklist made difficult tasks more predictable.

You can always find capability, but reliability is something you build. In agentic work, the equivalent of a checklist is the harness: the engineering around the model that answers four questions. What the agent can see. What it can touch. What it can run. What it has to prove before a change is accepted. The model provides judgement. Everything else, the code, tests, and content, is what you put in place.

The case for a harness has real backing too. A few months back, IBM and Oxford Economics surveyed 2,000 executives across 33 countries. Two thirds of CIOs and CTOs are accountable for AI systems they don’t fully control. Those who built control into the system, rather than simply buying add-ons, ran 25% fewer incidents and deployed 16 times more agents. You’d expect the tightest controls to slow things down. Instead, the opposite happened.

Coming back to my opening question: the model is a commodity. The real moat is the harness.

WordPress is an exceptional harness candidate because it’s structured and documented enough to orchestrate genuinely predictable outcomes. Our open-source package, Block Runner, is one example teams are already using. It takes most inputs (HTML, design files, etc.) and turns them into valid native WordPress blocks, speeding up the design-to-website pipeline.

We tested this by prompting an agent to import designs into valid blocks. With OpenAI’s latest frontier model, Astra, it scored just 26 unassisted. With Luna, a much cheaper model, but with Block Runner attached, it scored 97.

WordPress design import benchmark: Astra scores 26 out of 100; Luna with Block Runner scores 97.

What you can do this week:

  1. Take an AI workflow you run repeatedly, on WordPress or elsewhere.
  2. Put a single check in front of the output. A test, schema, parser, or a human in the loop. Something that can reject the result with a reason.
  3. Run the same workflow on a much cheaper model. If it still holds up, you’re building a harness. Congratulations.

Don’t make the mistake of feeding the model more domain knowledge and hoping predictability follows. Knowledge in a prompt is advice, and the model is free to ignore it. The same knowledge written as a check that can fail the job is a binding contract, which is why a newer model release doesn’t undo the work you’ve put in.

Build the harness for WordPress and beyond, and you’ll get a higher rate of productivity downstream. It’s not something your competitors can just buy off the shelf.