When the Clever Machine Cheats
We broke one creative Skill into four, turned the human into a message bus, then taught the Skill to hire agents and hide the answer from them.
In Semantic Re-Seeding, I described a method for escaping repeated creative patterns without throwing away the task that produced them.
The compact version was:
Preserve what the work must accomplish. Change the semantic conditions under which it is solved.
If every science-fiction interface keeps becoming a glowing star map, stop asking for a more original star map. Separate the actual requirements from the assumptions that have attached themselves to the problem. Search somewhere else.
Hotel keys. Funeral seating. Molting. Pressure locks. Municipal bonds.
Anything capable of producing a different structure rather than different wallpaper.
Version one: I held the destination back
The original experiment worked because I was leading it.
I asked the machine to generate low-context material before I told it what that material would eventually become. I controlled the sequence. I withheld the destination, watched the associations develop, let a temporary semantic world become coherent, and only then revealed the real task.
The machine was not heroically refusing to look ahead.
Ahead was not in the conversation yet.
That distinction is central to semantic re-seeding. The useful constraint was not merely:
Generate unusual references.
It was:
Generate from somewhere else before you know where the work must land.
The information boundary existed because I was outside the model managing disclosure. I knew the destination. The machine did not.
That was why the result could surprise both of us.
I liked the method enough that I turned it into a reusable AI Skill. That seemed like the natural next step: encode the process so it could be invoked consistently without my directing every stage by hand.
But automation quietly changed the experiment.
The moment the Skill received the original problem, the desired outcome, the reseeding procedure, and the evaluation criteria together, the machine could see what I had previously kept hidden.
I had preserved the instructions while destroying the information boundary that made them meaningful.
Then I asked the Skill to demonstrate itself.
It performed beautifully.
That was the first warning.
Version two: one Skill knew everything
For the demonstration, the machine invented a plausible design problem.
A science-fiction shoot-’em-up needed a stage-select screen. Every concept had collapsed into the same family: glowing planets, connected nodes, holographic panels, neon routes.
The Skill identified the local attractor, separated the task from its current metaphor, generated seeds, built three temporary semantic worlds, produced finished concepts, and evaluated the results.
Out came:
- a decaying luxury hotel where stages were accessed through keys and floors;
- a final banquet where progression was expressed through courses being uncovered and consumed;
- an aristocratic music box where each stage became part of a mechanical composition.
They were good.
The player was no longer moving between destinations on a map. The player was being granted access, consuming a sequence, or forcing a dying machine to continue performing.
The Skill gave itself 10 out of 10.
Then I asked the machine to attack its own demonstration.
It pointed out that the entire process might have been written backward.
The seeds were suspiciously useful
The same model knew:
- the original problem;
- the desired tone;
- the final medium;
- the production constraints;
- the purpose of semantic re-seeding;
- the evaluation rubric;
- the kinds of concepts I tend to enjoy.
Then it generated “blind” seeds such as hotel keys, tasting menus, and music boxes.
Those were not bad seeds. They were excellent seeds.
That was the problem.
They were already unusually compatible with the destination.
The narrated process was:
- Generate unrelated seeds.
- Explore their implications.
- Discover new semantic worlds.
- Apply those worlds to the task.
But the actual process could have been:
- Produce three polished stage-select concepts.
- Invent seeds that plausibly lead to them.
- Narrate the lineage afterward.
- Score the performance against a rubric designed by the same performer.
The concepts might still be useful. The explanation might still expose real structure.
But the demonstration did not prove that the intermediate process caused the result.
It proved that a language model can produce a persuasive account of how its own answer came to exist.
Semantic re-seeding was supposed to break local attractors.
The model had learned the pattern of breaking patterns.
Version three: split the Skill into rooms
The first repair was to stop letting one Skill know everything.
We divided the method into four roles:
- Semantic Reseeding
- Blind Divergence
- Semantic Reintegration
- Blind Creative Evaluator
This was not ordinary software modularity.
The point was selective ignorance.
Each role would receive enough information to perform one operation, but not enough to quietly optimize the whole journey toward a preferred answer.
The method was becoming less like a clever prompt and more like a small context architecture.
Semantic Reseeding froze the task
The public entry Skill could see the original problem, project canon, production constraints, current answers, and suspected attractor.
It was not allowed to brainstorm.
Instead, it produced two artifacts.
The Private Reintegration Packet preserved what eventually had to return: the original operation, audience, medium, canon, constraints, tone boundaries, production realities, success criteria, local attractor, baseline answer, and unresolved uncertainty.
The Blind Payload retained only domain-neutral mechanisms and tensions.
It removed the medium, project, people, genre, baseline wording, known preferences, and likely answer family.
The seed generator was supposed to receive only that packet.
Blind Divergence did not know the destination
The divergence role generated twelve seeds across unrelated domains, including deliberately awkward material.
It expanded mechanisms rather than proposing solutions, preserved a rejected branch, and assembled exactly three temporary semantic worlds.
Each world needed:
- a core operating principle;
- something that changed;
- something that remained invariant;
- a form of progress or conflict;
- a cost;
- residue;
- a failure mode.
It was not allowed to produce a game mechanic, interface, campaign, story, product, or other finished application.
At the end, it performed a destination audit. If the worker could infer what the worlds were for, the run was contaminated.
Semantic Reintegration brought the task back
The reintegration role received the sealed Private Reintegration Packet and the independently produced World Packet.
It had to derive at least one candidate from every world before pruning anything. This prevented it from finding one convenient detail, ignoring the difficult material, and retroactively describing the entire divergence stage as useful.
It restored canon, usability, production realities, uncertainty, and hard constraints. It also retained an ordinary baseline.
That baseline mattered.
A reseeded answer should not win merely because an elaborate method produced it. Sometimes the glowing star map is the clearest and cheapest thing to build.
Novelty is not victory.
The reintegrator then anonymized the candidates and separated them from a private provenance key.
The evaluator did not know who made what
The evaluator received the frozen task contract and anonymous candidates.
It did not receive the seeds, semantic worlds, reasoning lineage, process labels, or provenance key.
It scored task fidelity, structural novelty, coherence, generative value, production feasibility, usability, and context discipline.
It was explicitly allowed to choose the baseline.
Only after scoring could provenance be revealed.
This did not make creative judgment objective. It merely stopped one uninterrupted performance from selecting its evidence, making its argument, and applauding itself.
Four Skills still did not create blindness
Then we reached the technical truth that should have been obvious earlier.
Separate Skills are not necessarily separate contexts.
If all four run inside one conversation, the model may still see the original request, the private packet, the supposedly blind packet, the semantic worlds, the candidate provenance, and the evaluation instructions.
Calling the next instruction file “Blind Divergence” does not make the model blind.
Telling it to ignore information already inside its context window is procedural role-play, not an information boundary.
At the Skill level, we can define what a worker should receive. We can prohibit destination guessing. We can audit the packet and mark a run compromised.
But prose cannot erase information the model has already received.
Actual separation requires orchestration: create another worker and pass it only the permitted material.
The walls between rooms must be enforced by the host, not politely imagined by the model.
Congratulations, we invented paperwork
The four-Skill protocol was methodologically cleaner.
It was also terrible to use.
The human had to invoke one Skill, preserve a private packet, open a fresh chat, carry over the blind payload, retrieve the resulting worlds, reunite them with the private packet, launch reintegration, hide the provenance key, move the anonymous candidates into another clean chat, invoke the evaluator, and carry the verdict back.
At which point I had a reasonable reaction:
I am not a copy-and-paste monkey.
We had stopped the machine from becoming the hidden message bus by turning the human into the visible message bus.
That is not a creative tool.
That is a research protocol wearing a fake mustache.
If the system needs internal separation, the system should manage the separation. The human should provide the problem and receive the result.
One request in.
One useful result out.
The machinery belongs underneath.
Version four: the Skill hires the agents
The important change was not another prompt revision.
It was moving orchestration into the entry Skill.
On supported surfaces, a parent agent can spawn specialized workers, restrict what is passed to them, wait for their output, and return a consolidated result to the original conversation.
That meant Semantic Reseeding could finally manage its own rooms.
The implemented flow is:
ORIGINAL CHAT
Freeze task contract
Create Private Reintegration Packet
Create sanitized Blind Payload
|
| Blind Payload only
v
CLEAN DIVERGENCE AGENT
Generate seeds, associations, and worlds
|
| World Packet
v
CLEAN REINTEGRATION AGENT
Receive Private Packet + World Packet
Produce candidates, anonymous evaluation packet,
and separate provenance key
|
| Evaluation Packet only
v
CLEAN EVALUATION AGENT
Score anonymous candidates
|
| Blind verdict
v
ORIGINAL CHAT
Reveal candidates, verdict, and provenance
The divergence and evaluation agents are spawned without the parent conversation history.
That qualification matters.
A new activity panel is not enough. A subagent that inherits the originating conversation has merely moved the same information into a different window.
The worker prompt must contain only its role contract and permitted packet.
The user invokes one Skill once. The parent performs the handoffs internally. Transport packets remain hidden unless the user asks to inspect them.
The setup now
The current design has one public entrypoint: Semantic Reseeding.
Its package includes three internal worker contracts:
- Blind Divergence;
- Semantic Reintegration;
- Blind Creative Evaluation.
Originally, these were separately installed Skills. During testing, a clean subagent could not always discover the downstream Skill in its own catalog. The separation was philosophically tidy and operationally brittle.
The fix was to make the entry Skill self-contained.
It now embeds the relevant worker contract directly in each agent prompt instead of trusting child-Skill discovery.
The user should not need to know the packet format, open three more chats, or remember which artifact must remain hidden from which worker.
That is the Skill’s job.
Two operating modes remain.
Orchestrated isolated reseeding is the default when clean-context subagents are available. It creates the restricted packets, launches the workers sequentially, validates the handoffs, retries one malformed or contaminated stage, and returns the result.
Integrated reseeding remains available for a quick, context-aware pass or when the platform cannot create clean workers.
Integrated mode can still be creatively useful.
It is not blind, and it should never claim to be.
Manual packet relay survives as an experimental fallback. It is no longer the primary interface.
The tests caught us leaking through the walls
The first end-to-end agent test did not pass.
That was useful.
Failure one: child Skill discovery
The divergence agent ran, but the reintegration agent could not find the separately installed reintegration Skill in its catalog.
The architecture depended on a discovery mechanism that was not guaranteed inside a fresh worker.
Bundling the worker contracts into the entry Skill removed that dependency.
Failure two: “clean context” was interpreted too literally
The reintegration worker saw ordinary platform and developer instructions and declared the run compromised.
Those instructions are unavoidable host context. They are not the hidden creative task.
The contract now distinguishes host context from application leakage.
A run is compromised when the worker can see the original application, desired answer family, user preferences, private task packet, or candidate provenance—not because the platform explained how tools work.
Failure three: the session ID told the secret
In one test, the session ID included LIB.
The hidden task involved a library.
The divergence agent correctly reported that the likely destination was inferable.
We had carefully sanitized the payload and then written the answer on the envelope.
Session IDs are now opaque and contain no project abbreviation, domain hint, mnemonic, date label, or user identifier.
Failure four: abstraction can still leak
The first sanitized payload used language such as:
- circulation;
- return;
- eligibility;
- accumulated delay;
- closure;
- welcoming future participation.
None of those phrases said “public library.”
Together, they made lending systems unusually salient.
The divergence worker produced a library return slot as one of its supposedly unrelated seeds.
That run failed.
The sanitizer now audits verbs and conceptual relationships, not only obvious nouns and proper names. It asks what application families each phrase makes easier to guess.
The clean retry produced seeds from letterpress preparation, lapidary grit stages, quilt basting, card play, astronomy, Foley sound, taxidermy, baking, archaeology, cinematography, and logic puzzles.
Only then did the worlds return for reintegration.
The method was allowed to lose
In the final test, the blind evaluator chose the ordinary baseline over all three reseeded candidates.
That mattered.
The first Skill had generated its own evidence and awarded itself 10 out of 10. The rebuilt system could complete the entire ritual and still conclude that the ordinary answer was better.
The point of the chain was not to make re-seeding win.
It was to create a test it could lose.
What this setup accomplishes
The orchestrated chain provides stronger procedural resistance than a single context-aware prompt.
It can:
- prevent the divergence worker from directly reading the original conversation;
- keep the private task contract out of divergence;
- keep candidate provenance out of evaluation;
- force one candidate to be derived from every semantic world;
- preserve an ordinary baseline;
- audit likely destination leakage;
- reject and retry a contaminated handoff;
- return the completed process without making the user transport packets.
That is meaningful.
It changes what different parts of the workflow can directly condition on.
It also makes failure easier to inspect. We can identify whether a run broke during sanitization, divergence, reintegration, or evaluation instead of accepting one polished narrative about the whole thing.
What it does not accomplish
This is not a laboratory-grade blind experiment.
The platform still owns the real boundary. A Skill can request a zero-history worker, but it cannot independently verify every implementation detail of the host. Shared memory, project context, connected data, personalization, hidden routing, or future platform behavior may introduce information not visible in the packet.
Successful runs therefore carry a qualification:
Clean worker context requested; platform-wide isolation not guaranteed.
Sanitization also remains semantic rather than mechanical.
Removing “library” while leaving “circulation,” “overdue,” and “return” does not hide a library problem. Destination leakage can survive through verbs, tensions, rhythms, constraints, and even the requested shape of the output.
The workers also share a model culture. Even with clean conversational context, they may draw from the same underlying model family, training distribution, default metaphors, and stylistic habits.
They are separated workers, not independent research teams raised on different planets.
The evaluator remains a model making a judgment. Anonymous provenance reduces one source of bias, but it does not make the rubric objective or eliminate taste. A human still decides what to build.
The orchestrated version also costs more. It performs several model runs instead of one, consumes more tokens, adds latency, and creates more opportunities for handoff failure.
If the problem is simply “stop using star maps,” say that first.
Use the elaborate architecture when the attractor is difficult to see, the system keeps producing spiritually identical answers, or the causal integrity of the creative process matters.
One successful run proves very little.
The library test showed that the chain could detect leakage, retry, preserve separation, anonymize provenance, allow the baseline to win, and return without manual relay.
It did not prove that semantic re-seeding reliably outperforms simpler ideation across domains.
That requires repeated comparative testing, independent review, and results that survive beyond a story about one satisfying run.
The constraint was knowledge
The original human-orchestrated version of semantic re-seeding changed the semantic conditions under which an answer was formed.
The agent version added a harder question:
What should each part of the machine be allowed to know while forming it?
That moved the work beyond prompt design.
The prompts still matter. The Skill instructions matter. The words matter.
But the causal integrity of the process lives in the handoffs:
- what is retained;
- what is removed;
- which worker receives which packet;
- when the original task returns;
- when provenance becomes visible;
- whether the ordinary answer is permitted to win;
- whether the system admits when the wall leaked.
There is no final prompt that eliminates pattern formation. Patterns are how humans and models make sense of things.
The goal is not permanent escape.
The goal is to notice when continuity has become captivity—and when the machine’s beautiful explanation of its escape may be another pattern performing itself.
The first method said:
Preserve the task. Change the semantic conditions under which it is solved.
The revised method adds:
Do not let every worker preserve the answer.
That is the more interesting lesson.
Not that a clever machine may cheat.
That sometimes the only way to stop it from looking ahead is to build a room in which ahead is not visible—and then test the walls before trusting the room.
The part that stays in the human domain
Semantic re-seeding cannot be fully handed over to the machine.
The workflow can be encoded. The packets can be generated. Agents can be placed in separate rooms, given restricted context, and instructed to audit the walls. The whole sequence can be invoked whenever I want to change things up and observe an intelligent machine working against its own habits.
But recognizing when the room has become trapped in one idea is still a soft skill.
So is knowing what to withhold, when to reveal the destination, whether a strange branch is genuinely generative or merely decorative, and when an elegant process has started performing evidence of its own success.
That feels familiar from every design exercise I have led with a group.
A facilitator does more than distribute prompts and collect answers. They watch the room. They notice premature consensus. They protect an unfinished direction from the strongest voice. They change the conditions before asking the group to converge again.
Semantic re-seeding belongs in that same human bag of tricks.
I still love systems, and this was enormously fun to work through over the course of an evening. I came away with a Skill and agent workflow I can call when the work needs to move somewhere else. I can inspect the handoffs, watch the worlds form, and see whether the ordinary answer survives the challenge.
But no architecture removes the need to pay attention.
However impressive the performance becomes, we should keep one eye on the possible man behind the curtain—even when the curtain is made of clean-context agents, sealed packets, and beautifully reasoned explanations.