Raw dump
One raw thought dump, three sessions. The Codex line and the Claude line each ran the full pipeline — raw notes, reflection, voice-linted copy, Assad's words only, LinkedIn. The Quillbot line made one stop. Follow any line; they all terminate at what we learned.
Codex line · stop 1
Raw notes
I use a command to have the agent confirm understanding and give my thoughts back in an organized fashion. Start with my jumbled thoughts, then see how Codex converted my thought blob into organized thoughts.
- Codex is fast, token-light, concise, and linear.
- It works well when moving straight toward an outcome.
- It does not rewind well.
- Coding and artifact creation fill its context window faster.
- It is less useful when the work involves deviations or side paths.
- Claude Code tends to revert to a generic voice.
- There is no recorded data yet, but the engineering community seems to agree that Claude uses more tokens.
- Claude often uses more words to say the same thing and punctuation that people do not naturally use.
- Claude’s rewind works very well.
- /btw becomes a killer command once you understand it: ask an inline question, clear that question, then continue along the original context.
- Claude can hold more in its context window. That may relate to its heavier token use.
- Semantic planning is producing more consistent results.
- Yesterday: delivered the Next.js Better Hoops demo explaining the short version of the semantic layer.
- The day before: built the operations semantic layer and saw an immediate improvement in usefulness.
- Yesterday: updated security. Details are confidential.
- Used tooling to draft 10 specific emails for different prospects and clients in parallel.
- Turned a backlog into one orchestrated task using the Warp terminal.
- Last week: built the analytics page behind a web application for friends looking for jobs.
- The business is taking a different approach: everything published will be written with AI.
- One version will come from Codex. Another will come from Claude.
- The open question: does the difference between them matter?
Codex line · stop 2
Reflection
The command also gives back the agent's synthesis in the form of a reflection — what the agent understands. At this point I can start over or make changes. We don't want to have a conversation. We want to get started on the right path with any agent and work in a linear fashion, so we want to know how the agent understands our thoughts.
The core idea I hear is: Codex and Claude have meaningfully different working styles, but the real test is not which model sounds better in theory. It is whether those differences change the usefulness, quality, or effect of the work when both are used to produce real business output.
Is that the idea, or is the sharper point that the business will publicly test whether customers can tell—or care—which AI wrote each version?
Codex line · stop 3
Voice-linted output
Next the agent runs through a voice-specific set of instructions. This is where we add our authentic voice — the way we ask the machine to communicate on our behalf, the language we try to avoid, and the way we strive to be when we communicate.
As a reader who uses AI for real work, I need this to explain how each model changes the process and why the finished result is the real test.
Working copy:
Codex and Claude create different working conditions. We are testing whether that changes the finished work.
Codex is fast, concise, and linear. Its context fills quickly during coding and artifact creation. Claude uses more words, but holds more context and makes rewind useful. Once learned, /btw changes the workflow.
This week, those workflows helped us ship the Better Hoops demo, build an operations semantic layer, and draft 10 emails in parallel.
Now /prznt will publish two versions of its writing: one from Codex and one from Claude.
Does the difference survive into the finished work? Does the reader care?
The linter only flags “Better” in the proper name “Better Hoops.” That is a false positive.
Codex line · stop 4
Assad's words only
This was a deviation. Instead of asking the agent to interpret my voice, the instructions were: use only my words.
Codex is fast. Token light. Concise. Keep it linear.
Coding and artifact creation fill the context window faster. Not good for deviations. Good for going straight. Codex doesn’t rewind well.
Claude Code is token heavy. More words to say the same. Uses punctuation that humans don’t. Reverts to the generic voice.
Rewind works great. /btw is the killer command once you learn to work with it. Ask a question inline, clear the question, and continue the linear context window. Claude can also keep more in the context window. Makes sense because it uses more tokens.
Semantic planning. Consistent results.
Yesterday we delivered the Next.js Better Hoops demo to explain the short version of the semantic layer. The day before, we built the semantic layer for the operations layer, seeing an immediate change in usefulness. Security updated yesterday. Details confidential.
Used tooling to draft 10 specific emails to different prospects and clients in parallel. Turned backlog of work into one orchestrated task using the #warp terminal. Last week, built the analytics page behind a web application for job hunting friends.
As a business, we’re taking an entirely different approach. Everything we write is done with AI. One version will come from Codex. One will come from Claude.
Does the difference matter?
The linter only flags “Better” in the proper name “Better Hoops.” No added language appears in the copy.
Codex line · stop 5
The next command for both agents was to write for a LinkedIn post.
Codex is fast. Token light. Concise. Good for going straight. Not good for deviations.
Claude Code is token heavy. More words to say the same. Reverts to the generic voice. Rewind works great. /btw is the killer command once you learn to work with it.
Yesterday we delivered the Next.js Better Hoops demo. The day before, we built the semantic layer for the operations layer, seeing an immediate change in usefulness.
As a business, we’re taking a different approach. Everything we write is done with AI. One version will come from Codex. One from Claude.
Does the difference matter?
Claude line · stop 1
Raw notes
Claude was given the same instructions. See the difference.
Codex
- Fast, token-light, concise. Keeps it linear.
- Doesn't rewind well. Coding and artifact creation fill the context window faster.
- Not good for deviations. Good for going straight.
Claude Code
- Reverts to the generic voice. No data recorded, but engineering community agrees it's token-heavy — more words to say the same. Uses punctuation humans don't.
- Rewind works great. /btw is the killer command once you learn to work with it — ask a question inline, clear it, continue the linear context window.
- Can keep more in the context window — makes sense, it uses more tokens.
- Semantic planning yielding consistent results.
Recent work (proof points)
- Yesterday: delivered the Next.js Better Hoops demo — the short version of the semantic layer explained.
- Day before: built the semantic layer for the operations layer — immediate change in usefulness.
- Security updated yesterday; details confidential.
- Tooling drafted 10 specific emails to different prospects/clients in parallel — backlog turned into one orchestrated task using the Warp terminal.
- Last week: built the analytics page behind a web application for job-hunting friends.
The closing move
- As a business, we're taking an entirely different approach: everything we write is done with AI.
- One version comes from Codex, one from Claude. Does the difference matter?
Claude line · stop 2
Reflection
The command also gives back the agent's synthesis in the form of a reflection — what the agent understands. At this point I can start over or make changes. We don't want to have a conversation. We want to get started on the right path with any agent and work in a linear fashion, so we want to know how the agent understands our thoughts.
The core idea I hear is: the two tools have real, observable personalities — one linear and lean, one verbose but able to hold more and rewind — and instead of picking a winner, you're running both and putting the question to the reader: when everything is AI-written, can you even tell, and does it matter? The recent shipped work is the credibility floor under the comparison — this isn't a spectator's take, it's from daily production use.
Is that the right read? And is the target a LinkedIn/newsletter-style post, or something else?
Claude line · stop 3
Voice-linted output
Next the agent runs through a voice-specific set of instructions. This is where we add our authentic voice — the way we ask the machine to communicate on our behalf, the language we try to avoid, and the way we strive to be when we communicate.
Two tools write for this business now. Codex is fast, token light, linear. It goes straight, and it fills its context window quickly. Claude uses more words to say the same thing, holds more in context, and rewinds cleanly. Same work, two personalities.
The pairing shipped real things this month. A Next.js demo that explains the semantic layer in short form. A semantic layer under our own operations. Ten specific prospect emails drafted in parallel from one orchestrated task. An analytics page for friends on the job hunt.
Everything we write is done with AI. One version comes from Codex, one from Claude. Does the difference matter?
Lint clean, rubric clear. One call to flag: the client name “Better Hoops” trips the forbidden-word check (“better”), so the demo is unnamed in the draft. Say the word and it goes back in. Left out: /btw, rewind mechanics, and the security update, since the punchy version carries the point; a longer teardown post could hold the tool details if you want one.
Claude line · stop 4
Assad's words only
This was a deviation. Instead of asking the agent to interpret my voice, the instructions were: use only my words.
Codex is fast. Token light. Concise. Keeps it linear. Doesn't rewind well. Coding and artifact creation fill the context window faster. Not good for deviations. Good for going straight.
Claude is token heavy. More words to say the same. Uses punctuation that humans don't. Rewind works great. Claude can also keep more in the context window. Makes sense because it uses more tokens. Semantic planning yielding consistent results.
Yesterday we delivered the Next.js Better Hoops demo to explain the short version of the semantic layer. Day before, built the semantic layer for the operations layer and seeing an immediate change in usefulness. Drafted 10 specific emails to different prospects and clients in parallel. Turned a backlog of work into one orchestrated task. Last week, built the analytics page behind a web application for job hunting friends.
As a business, we're taking an entirely different approach. Everything we write is done with AI. One version will come from Codex, one will come from Claude. Does the difference matter?
This is your wording rearranged, nothing added beyond connective trims. Note it now carries terms the voice archive forbids in drafted copy (“great”, “entirely”, em-dash-free at least) since your raw words override the lint. Dropped: /btw and the Warp terminal mention, and “Claude Code reverts to the generic voice” since the post is written by Claude. Say the word to put any back.
Claude line · stop 5
The next command for both agents was to write for a LinkedIn post.
Everything we write is done with AI. One version comes from Codex, one from Claude.
Running both daily, the personalities are clear.
Codex is fast. Token light. Concise. Keeps it linear. Good for going straight, not good for deviations. Coding and artifact creation fill the context window faster.
Claude is token heavy. More words to say the same. But rewind works great, and it keeps more in the context window. Makes sense, it uses more tokens. Semantic planning is yielding consistent results.
What the pairing shipped recently:
- A Next.js demo for Better Hoops explaining the short version of the semantic layer
- A semantic layer under our own operations, with an immediate change in usefulness
- 10 specific emails to different prospects and clients, drafted in parallel from one orchestrated task
- The analytics page behind a web application for job hunting friends
As a business, we're taking an entirely different approach.
Does the difference matter?
Kept your words, reshaped for LinkedIn scannability: hook moved to the top, ship list as bullets, question as the closer. Left out /btw, Warp, and the security update to keep it tight. Want any of those back in?
Quillbot line · only stop
Social media output, first pass
Writing for LinkedIn is an option at Quillbot. I gave Quillbot the full word dump and tested the free option too.
Ever wonder how AI tools shape the way we work and create?
We’ve been exploring two very different AI companions: Codex and Claude. Codex is fast, concise, and excels when you need to go straight to the point—though it doesn’t handle detours well. Claude, on the other hand, uses more tokens, speaks more like a human, and shines when you want to rewind or explore ideas in a flexible flow. It even has a killer command, /btw, that lets you ask questions inline without breaking the thread.
Just yesterday, we launched a demo for a semantic layer that’s already changing how we build. We’re drafting emails, orchestrating tasks, and creating tools—all powered by AI. And here’s the real question: does it matter which AI we use? Both bring something unique, and together, they’re helping us rethink how work gets done.
What’s your experience been with AI tools? Do you prefer speed and focus, or flexibility and depth?
Terminus · all lines
What we learned
What we think after testing. We thought the best writer was Quillbot on the first pass; after a second and third, we decided it was being presumptive, and the goal in our writing is to present facts. Of the three, the Codex agent harness paired with GPT 5.6 Sol produces the output we trust most. We find that Claude with Fable 5 tries too hard — it's writing the narrative for the presentation. We're ending the exercise deciding that we are fine with content published by an agent. Our goal is to get the point across and back up our statements. Today we choose not to prioritize quality of prose.
- The process we use when working with agents is often: raw thoughts, confirmation of understanding, refined thoughts, confirmation, continue.
- Unfortunately, organized raw thoughts do not make a narrative.
- Of the three narratives, Quillbot's was the one we enjoyed reading most. All three came from the same raw thought dump.
- Designing the tool for a consistent agentic voice that we trust and like is a high priority.
- Short statements are hard. Complex thinking is so far impossible.
- Raw thoughts need to be organized into a narrative with a beginning, middle, and end.
- Frontier developers will figure this out. Today none of the models or agents are 100%. But still gamechanging tools.