I audited nearly 3 billion Codex tokens. Astra won the hard tasks. Sol was the better deal

A dramatic image depicting a battle between two characters, Astra on the left representing cool blue tones, and Sol on the right symbolizing fiery orange hues, with 'Astra vs Sol' text prominently displayed in the center.

Astra was burning through my Codex credits, so I had Codex audit its own local history.

I wanted a straightforward answer: was the extra cost buying better work? After 1,022 task turns and nearly 3 billion processed tokens, the honest answer was “sometimes.”

That “sometimes” changed both my configuration and the way I assign work to coding agents.

What I measured

I analyzed real development sessions from two TypeScript codebases. The sample included:

  • 1,022 token-bearing task turns
  • 27,487 model responses
  • 2.99 billion processed tokens
  • 188 attributable Git commits

I compared GPT-6 Astra at medium reasoning with GPT-5.6 Sol at high reasoning. These were the levels I actually used for each model, not matched experimental settings. The work included full-stack features, migrations, complex interface state, concurrency, security controls, accessibility, tests, and repository maintenance.

Nearly 3 billion sounds absurd because it is not 3 billion unique tokens. The total includes cached input and context processed again across responses. It is still a useful record of what Codex consumed while doing real work.

The comparison is messy in the way real project data is messy. The models did not receive matched prompts, and they worked during different stages of development. Treat this as a field study, not a laboratory benchmark.

The result I did not expect

Given the credit burn, I assumed Astra would be the more verbose model. It wasn’t.

Per response, Astra medium used an average of 101,128 total tokens. Sol high used 112,678. Astra also produced 275 output tokens per response versus 315 for Sol. Its reasoning-token average was 61, less than half of Sol’s 124.

In percentage terms, Astra used 10.3% fewer total tokens, 12.7% fewer output tokens, and 50.8% fewer reasoning tokens. Even its final reports were about 30% shorter.

So why did it feel so expensive?

At the time of this analysis, OpenAI’s ChatGPT credit rates (https://learn.chatgpt.com/docs/pricing) priced every Astra token category at 2.5 times the equivalent Sol rate. Astra’s lower token use softened that multiplier, but did not cancel it.

Sol’s current pricing is promotional through at least November 21, 2026; if that gap narrows, the economic conclusion changes.

After I normalized both models to Standard execution, the median estimated cost of one task turn was:

  • Astra medium: 31.8 credits
  • Sol high: 17.5 credits

Sol was about 45% cheaper. Viewed from the other side, Astra cost 82% more per turn.

The historical sessions were more expensive because many ran in Fast mode. Their estimated medians were 77.5 credits for Astra and 39.2 for Sol. OpenAI currently charges 2.5 times the Standard rate for Fast execution (https://learn.chatgpt.com/docs/agent-configuration/speed) on these models.

The starkest number was the aggregate: Astra accumulated about 13% more estimated credits across just 384 token-bearing turns than Sol did across 638.

Configuration helps, but pricing still wins the argument.

I read the diffs before judging quality

Token counts can tell me what a model consumed. They cannot tell me whether I would trust the result in production.

Astra’s advantage became visible when one task touched several parts of the system. In a filtering feature, it pulled the core logic into a reusable unit and handled unavailable assignees, UTC due dates, linked tasks, keyboard focus, accessibility, unit tests, and end-to-end coverage. In security and runtime work, it combined identity checks, deadline handling, proxy constraints, and invariant tests into coherent changes.

This was the kind of work where Astra earned its price. It was better at carrying a large mental model across files and noticing states that had not been spelled out in the prompt.

Sol was less ambitious, which often made it a very good engineering partner. It produced smaller, focused changes and was excellent at repository tooling. One example reused existing CI checks, kept command output bounded, rejected invalid arguments, and added the tests the tool needed. Sol also performed well during rapid UI iteration, where I could review a small change and immediately steer the next one.

The completion numbers were close. Astra completed 98 of 100 root tasks, while Sol completed 122 of 127. Every one of Astra’s 77 attributable commits included test changes. Sol did the same in 79 of 80.

Astra’s median attributed commit touched eight files and added 659 lines. Sol’s touched two files and added 181. I would not use that as a scoreboard though. It says more about the jobs I assigned than the models themselves, but it does explain why Astra looked stronger on broad work and Sol looked sharper on narrow work.

After reading the output rather than merely counting it, I gave Astra a small to moderate quality edge on the hardest autonomous tasks. Sol came close enough everywhere else that its price became decisive.

What I changed in Codex

Fast mode was the first thing to go. I turn it on only when getting an answer sooner is worth paying 2.5 times the Standard rate.

For Astra, I settled on medium reasoning, Standard speed, and low verbosity. Medium handled the harder work in this sample. There was no reason to pay for a higher reasoning level on every task.

I shortened my global instructions too. A useful AGENTS.md changes the agent’s decisions. Repeated reminders about planning, narration, and generic engineering habits only add context. I kept the security constraints and the checks each project actually requires.

I also removed a blanket TDD rule. The agent still has to prove its work, and the commit data shows that testing did not disappear. The difference is that a copy change no longer triggers the same ceremony as a concurrency fix.

My global skill list is smaller now. I load a skill when it adds knowledge or a tool that the task needs. Process-heavy skills stay out of ordinary work because Astra already plans, explores, implements, and verifies without needing several overlapping workflows in its prompt.

Finally, I stopped treating one conversation as a permanent home for a codebase. When the objective changes, I start a fresh task and carry over a short handoff: the goal, the relevant architecture, constraints, decisions already made, known failures, and the next concrete step. Old context is useful until it starts charging rent.

The setup I use today

For most work, my default is GPT-5.6 Sol with high reasoning, Standard speed, and low verbosity.

I move to GPT-6 Astra with medium reasoning when the task involves security, architecture, concurrency, a broad change across the codebase, or a long autonomous run where one missed state could be expensive.

If I wanted to use Astra all day, I would keep the same guardrails: Standard speed, medium reasoning, low verbosity, compact instructions, a narrow skill set, and a fresh task for each coherent feature.

How to audit your own usage

Do not trust the model name attached to the conversation. Codex can switch models during a task, so inspect the effective model and reasoning level turn by turn.

Separate Fast and Standard runs before calculating cost. Count cached input as well as uncached input, reasoning, and output. Then open the work itself. Read representative diffs, check whether tests exercise meaningful behavior, and look for follow-up fixes or abandoned changes.

Most of all, be honest about the sample. Real project history can answer, “Which model works better for me?” It cannot tell the whole internet which model will win on identical work.

My answer is now specific. Sol high is the better daily driver. Astra medium is the model I want when the task is hard enough that an overlooked edge case will cost more than the credits.

Leave a Reply

Discover more from Alex López

Subscribe now to keep reading and get access to the full archive.

Continue reading