The book was the point

I run my real operations with Claude as the primary AI. It has context, history, and a growing role in how I get things done. Then I brought aboard a second frontier model: OpenAI’s GPT-5.6, which I call Sol.

This is early. It is one person, one household, and a working experiment. There is no model-fusion product here yet. I am setting a precedent for how I want my operation to work before the habits harden.

The arrangement that emerged is simple: Sol does the work. Opus, the Claude model I use, judges it.

That sentence sounds more settled than the process was. I added Sol because two independent models give me a second instrument. One can notice an omission, weak assumption, or awkward choice that the other has stopped seeing. Competition makes that difference visible.

A real test: Olga’s book

My wife Olga wrote a style book in Russian. Translating it into English gave me a useful test because a good result depends on more than literal accuracy. The English needs to carry her ideas, rhythm, and presence. A smooth sentence can still be a bad translation if the author disappears inside it.

I had two models produce their own “best of both worlds” merge from the available English versions. Before either merge was judged, the models agreed on the rubric. That mattered: the standard could not be quietly rewritten to favor whichever draft looked stronger afterward.

One model then judged both versions blind. It evaluated the writing without knowing which model had produced which draft.

The winning version won on voice. It kept Olga sounding present and plain. The other version was polished, but some of that polish created distance. The judge preferred the translation that felt more like a person speaking clearly from the page.

That result gave me something useful. Then I nearly framed it the wrong way.

The comparison came back as: model X beat model Y. I pushed back. I had not run the exercise to manufacture a leaderboard. I wanted the best English version of Olga’s book. The models were instruments in that process; the book was the point.

So I reframed the result around what improved: the author’s voice survived the merge.

Why the split helps

A single model can draft, critique itself, and revise. I use that loop often. Its weakness is that the same model carries its preferences and blind spots through every pass. Self-critique can improve an answer while leaving the original frame untouched.

An independent judge changes the pressure. The worker has to satisfy an external reading of the rubric. The judge can catch choices that feel natural to the worker because they came from its own habits. Blind judging also removes the temptation to reward a model’s reputation instead of the artifact in front of me.

The value is not that one model is permanently the worker and another permanently the judge. Roles can change. Independence is the useful part: produce with one set of instincts, inspect with another, and keep the decision tied to an explicit standard.

This costs more. Two models plus a judging pass consume more time, tokens, and attention than asking one model to keep revising. The setup also creates new failure modes. Both models can agree on a poor rubric. A judge can sound decisive while missing the thing that matters. Blind evaluation reduces one bias; it does not create truth.

For important work, I still have to inspect the result and decide what ships. In this case, Olga’s book gives me another essential judge: Olga.

What I am keeping

I am keeping the worker-and-judge pattern for tasks where quality is subjective, mistakes hide inside fluent language, and the result matters enough to justify the extra pass. Translation is a strong fit. Strategy, research, and consequential writing may be as well.

I have one case, not a general proof. I do not yet know how reliably this pattern transfers across tasks, how often the added cost pays for itself, or when two models merely reinforce the same mistake. Those questions need repeated work with real artifacts.

The useful change is already clear in my own operation. I no longer treat a model’s polished answer as the natural end of the process. I can ask another capable system to examine the work from outside its frame. When they disagree, the disagreement shows me where judgment is required.

That is what model fusion means in my operation today: independent attempts, explicit criteria, blind comparison where possible, and a human accountable for the final call.

The first lesson came from a family book. The technology mattered. Olga’s voice mattered more.