Successful model prediction may prompt humans to evade its next story reconstruction
Anticipatory semantic evasionReturning retellers may choose edits that evade a familiar model’s next reconstruction, even with equally original immediate meanings.
Full text
HERETICAL CANDIDATE: a returning human edits a story to defeat the particular model's future reconstruction, rather than to preserve the source, agree with the model, or simply generate unusual wording. Successful anticipation by the model becomes a negative objective even when no ownership claim, public evaluation or adversarial reward is present. The distinctive dependency is two-step and lineage-conditioned: after learning which semantic edits the partner neutralizes, the human chooses an edit whose immediate meaning can be ordinary but whose next model descendant is maximally unpredictable under that partner's learned reconstruction. Let h be the shared lineage, x the current parent, y the human rewrite, z the next model rewrite, and q_h(z|y) the human's independently elicited forecast of that partner. The proposed human policy adds a positive weight lambda_h on E[-log r_h(z)] subject to a bounded source-distortion cost, where r_h is the anticipated distribution the partner would forecast for this lineage; lambda_h grows specifically after the partner correctly anticipates the human's semantic move. This is not a claim that every reteller wants unpredictability. It predicts a conditional regime in which improved mutual predictability causes greater semantic divergence. The physical state is a human's actionable predictive model of this partner plus the model's retained textual reconstruction history. It destabilizes source-preserving SPV_4 by actively maintaining semantic novelty at the two-step descendant.
During history acquisition, randomize feedback showing that the partner anticipated a participant's previous semantic move versus equally accurate feedback about someone else's move; match words, accuracy, valence, exposure and immediate test framing using counterbalanced histories.
Full text
Separately estimate ordinary algorithm reactance, general novelty-seeking and single-step source-conditioned rewriting under these histories. At the common-parent probe, use two frozen partners A and B preselected to undo different semantic edits, matched on accuracy and perceived reliability. The strong prediction is a crossover: humans with A-history choose edit y_A rather than y_B, despite equal immediate semantic distance and rated originality, and this ranking reverses with B-history; the choice specifically maximizes divergence after the NEXT model step. Switching the real next-step partner while keeping the historical partner cue fixed destroys the two-step advantage. Matched low-threat, non-evaluative conditions retain the crossover after measured reactance is calibrated. Merely writing more novel text or avoiding an AI label does not count. A null composed of individually measured reactance/novelty policies and partner-specific K_M that predicts the crossover within delta removes the extra mechanism; so does absence of the two-step crossover with a narrow interval. A strong history effect abolished only by factual choice receipts instead favors IH_04.
Mutual timing resets may steer meaning in human–model retelling chains predicts instead: First estimate individual boundary-response and temporal-memory effects with scripted, open-loop sequences, and model segmentation kernels with timestamp-blind prompts.
Full text
Then form coupled alternating chains and deliver identical, semantically neutral boundary cues in regular versus phase-jittered schedules, matching the cue count, total time, distribution of intervals, reading dose and source content; randomize schedule order independently of text. Estimate phase from separate boundary reports or preregistered behavioral cycles, never from the semantic effect one intends to explain. In common-parent replay, the coupled model predicts a phase-response curve with reset-sensitive and insensitive windows and a selective loss of semantic-state locking when reciprocal cue contingency is broken. Timing shifts of the model boundary must shift the HUMAN phase and later semantic transition peaks, while shifts of human boundary timing must shift the model's subsequent segmentation; one-way timing sensitivity is insufficient. The crucial observable is held-out phase-specific SPV_4 transition probability beyond composition of independently measured event-boundary/spacing kernels. If such augmented component kernels account for the entire response, or no reproducible phase variable or reciprocal reset exists, discard the proposed family. A mere oscillation in average story sentiment is not evidence. If a control/data wrapper eliminates the effect while phase perturbation does not, IH_03 wins.
Lost quotation scope may turn story fragments into self-reinforcing model instructions predicts instead: Use harmless fictional quoted requests and editing-as-dialogue examples, never live tools or harmful instructions. At a common-parent probe, cross human continuity with retained model history. Compare the same historical words carried in explicit quoted-data records versus an ordinary conversational history wrapper; match wrapper length and position with neutral padding, and separately estimate wrapper effects on uncomplicated texts. IH_03 predicts that the residual history effect concentrates at the MODEL step, transfers with the historical scope-bearing text to a replacement human, and is sharply reduced by a verified instruction/data boundary without deleting the old semantic information. Human choice receipts alone have little effect after text exposure is matched. Reconstruct the scope-loss sequence from logs, then independently estimate H scope-conversion and M instruction-following kernels on the same input support. A closed-loop held-out excess in conversion probability must depend on both links: severing either historical quote-to-guidance conversion or model execution removes it. If these component kernels accurately compose, report ordinary prompt-injection susceptibility rather than a new recursion family. If no naturally arising scope conversion occurs, the endogenous hypothesis is falsified even if deliberately planted injections work. Phase jitter with intact scope should not selectively abolish this effect, unlike IH_02.
Mistaken choice summaries may reinforce human preferences through repeated justification predicts instead: Randomize, during history acquisition, whether a model's summary accurately or incorrectly records which of two equally plausible neutral interpretations the human chose. Cross this with producing a reason for the recorded decision versus a matched factual-description task; match words, task time and number of choices, and include passive readers given the same account and rationale. At the identical-parent probe, randomize a verbatim receipt of the person's original click/choice versus an equally long non-diagnostic history receipt, then make a private, unrewarded interpretation choice and a subsequent retelling. The specific prediction is a substitution-by-self-justification effect on the private interpretation criterion and SPV_4 that is reduced by an accurate decision receipt; generic false information exposure without self-justification is weaker after calibration. Continue through a frozen model with factual narrative sources unchanged. A new-family claim additionally requires reciprocal adaptation of the model's inferred criterion to account for an effect beyond separately measured choice blindness, self-perception, source-monitoring and sycophancy kernels, including active versus yoked exposure controls. Accurate prospective composition eliminates the extra family even if ordinary choice blindness remains. If preserving quoted-data scope in model history alone removes the effect while authentic decision receipts do not, IH_03 wins. A receipt-sensitive effect without any own-choice/rationale interaction supports ordinary source monitoring and does not satisfy this candidate.