What if a model's new skill leaked out through answers that have nothing to do with that skill? A new paper, Post-Training Leaves Behavioral Shadows on Unrelated Decisions, shows exactly that: a small model got better at coding without ever seeing a line of code, just by copying one word at a time from a model that had been trained to code.
- What happened: researchers moved part of a model's coding skill into a fresh copy using only one-word answers to unrelated prompts.
- The result: the student scored 51.22% on the HumanEval+ coding test, the same as its teacher, against 45.88% for a matched control: a gain of 5.34 percentage points.
- The data: just 5,664 prompt-and-word pairs, picked from 17,858 near-tie questions.
- It held up: across 15 repeated runs the gain averaged 4.80 points, and every run was positive.
- The limits: it only works when the teacher and student start from the same public model, and the gains are modest.
What did the researchers do?
The team, from Peking University, Georgia Tech, ShanghaiTech University, Tsinghua University and Lovart AI, started with a public model, Qwen2.5-1.5B-Instruct. They trained one copy privately to code better. That copy is the teacher. A second, untouched copy is the student.
The student never sees the teacher's code, its weights or its training data. It only sees which single word the teacher picks in answer to ordinary prompts that have nothing to do with programming. The authors call the method Active Taskless Distillation.
How can one word carry a skill?
Every time a language model writes, it gives each possible next word a probability. You can see this step in our lesson From logits to tokens.
The trick is to find near-ties: prompts where the original model is almost 50-50 between two everyday words. On a knife-edge like that, a tiny change inside the model is enough to tip the choice. Training the teacher to code changes its insides slightly, and those changes nudge which way it tips, even on questions about cooking or travel. The authors call this the behavioral shadow of training.
The student then learns from these prompt-and-word pairs alone. Thousands of tiny nudges add up, and some of the teacher's coding skill comes along with them.
Think of two people who learned to write from the same teacher. One of them later takes a coding course. Their handwriting doesn't change, but in thousands of tiny choices, like which of two words to use, small habits shift. Copy enough of those choices and you start to pick up a little of what they learned.
How big is the effect?
| Measure | Result |
|---|---|
| Training data | 5,664 prompt-and-word pairs |
| Student on HumanEval+ | 51.22% |
| Teacher on HumanEval+ | 51.22% |
| Matched control | 45.88% |
| Gain over control | +5.34 points (95% confidence interval 1.22 to 9.60) |
| Average over 15 runs | +4.80 points, all positive |
The effect also showed up on six multiple-choice tests, including ScienceQA, OpenBookQA and HellaSwag, with gains between 0.81 and 2.90 points.
Where does it fail?
The authors are clear that this is "an initial exploration":
- Same starting point needed. The transfer only worked when teacher and student came from the same public model. It failed across model families.
- Low bandwidth. When the teacher's own improvement was small, the student gained little. On the MBPP+ test, a teacher gain of 2.38 points became a student gain of only 0.33.
- Not everything travels. Memorised answers and hidden ciphers did not transfer at all.
Why does this matter?
First, it is a striking result about how models work. Training for one skill leaves faint fingerprints on unrelated behaviour, and those fingerprints carry real information.
Second, it raises a question for anyone who fine-tunes a public model and offers it as a service. If others can see your model's answers and know which public model you started from, they may be able to copy part of what you added. That is our reading of the result, not a claim the authors make, and the gains so far are small. To learn how post-training fits into building a model, see How a modern LLM is built.
Read on 3 October 2026. The paper is an arXiv preprint (submitted 24 September 2026), so other scientists have not yet reviewed it, and all results come from one small model family. Read the paper.