Alignment needs preference pairs. Fine-tuning needs examples. Distillation needs a corpus to run the teacher over. In almost every real project, that data does not exist in the quantity required — so generating it has become a standard step rather than a last resort. Interviewers ask here because it is where theory meets what teams actually do, and because the failure modes are subtle.
What gets asked
Why generate data at all. Four distinct reasons, and they call for different techniques: rare events you cannot ethically or practically collect more of; privacy restrictions that block the data you most want; class imbalance that lets a model score well by ignoring the minority; and speed, because real data arrives on its own schedule.
How it works in a language pipeline. A strong model produces training data for a weaker one. The established pattern starts from a couple of hundred hand-written seed tasks, prompts a model few-shot to invent more instructions of the same kind, has it produce inputs and responses, filters the results, and iterates. A small seed becomes a large instruction-tuning set.
Why filtering is the real step. Generation is cheap and produces duplicates, malformed items and wrong answers alongside good ones. Deduplication, format validation, closeness checks against the existing pool and automatic correctness checks where possible are what turn raw output into a dataset. Candidates who describe generation and skip filtering are describing the easy half.
The classical toolkit. Beyond language: sampling from distributions you specify (transparent, controllable, contains no structure you did not put there); interpolation between real minority-class examples, which assumes the line between two real points contains plausible points and fails for categorical features; learned generative models for tables, which capture non-linear multi-column structure at the cost of opacity; and augmentation for images and audio, where the label is already known.
Evaluating it. Three axes, because passing one proves nothing about the others. Realism — do the distributions match, field by field. Utility — train on synthetic, evaluate on held-out real data, which is the test that counts. Privacy — how close does each synthetic record sit to a real one.
The follow-ups that catch people
Is synthetic data automatically safe to publish? No, and this is the trap. A generative model can overfit and reproduce training records closely, and rare records — the most sensitive ones — are the likeliest to be memorised. Safety is measured with nearest-neighbour distance checks and membership-inference tests, not assumed from the fact that a model produced it.
What is mode collapse, and why is it especially bad here? Wanted: the generator learns to emit a few common patterns and drops the rest of the distribution. The irony is that people generate data precisely to strengthen the rare cases, and collapse removes exactly those.
Where does bias come from in synthetic data? Wanted: inherited from the source and often sharpened, since a model trained to produce likely examples reproduces the majority pattern more reliably than the data did. Auditing the generation process matters as much as auditing the trained model.
Your synthetic set passes every distribution test. Is it good? Wanted: those tests look at one field at a time, so correlations between fields can be destroyed while every margin matches. Utility testing on real data is what settles it.
Why build a very large model you cannot afford to serve? Wanted: as a teacher. Its outputs become training data for smaller deployable models, and that value outlives the cost of serving it.
How it gets worded
Generating data
- "You have a badly imbalanced classification problem. How would you generate data for it?"
- "When is interpolating between minority examples the wrong tool, and how would you choose between that and a learned tabular generator?"
- "How do generative models for tables cope with categorical columns?"
- "Why do adversarial generators struggle on small datasets?"
- "Show me how you would decide whether a synthetic dataset is safe to hand to another team."
- "Could synthetic records leak something about real people? How would you detect it?"
- "How can generating data make an existing bias worse rather than better?"
- "What are the risks of using synthetic data in a regulated industry?"
Scale, data and compute
- "What do the scaling results tell you, and what do they not?"
- "How does performance move as you add parameters? What happens if you grow the model without growing the data?"
- "What is compute-optimal training, and why did a smaller model trained longer beat much larger ones?"
- "Fixed compute budget. How do you split it between model size and training tokens?"
- "How would you predict a larger model's score before you train it?"
- "Where do the returns start to flatten, and how does any of this feed into an architecture decision?"
Scaling laws and compute-optimal training do not have a lesson here yet. The questions above are included because they are asked often; for now, read the Kaplan and Chinchilla papers directly. The one line to carry into an interview: for a fixed compute budget, parameters and training tokens should grow together, and the models released before that result was published were badly undertrained for their size.
Reading path
- Synthetic Data Generation — the whole chapter: why, the generation loop, the classical toolkit, the three-axis evaluation, and the failure modes.
- Pruning and Distillation — distillation is the same idea viewed from the compression side, and it is where the data bottleneck bites hardest.
- Model Alignment — what preference data is and why collecting it is the expensive part.
- Noise, Augmentation and Parameter Sharing — augmentation as a regulariser, the classical root of the idea.
The framing that works
Synthetic data is best described as moving a constraint, not removing one. It converts a data problem into a modelling problem: you no longer need to collect examples, but you now need to be right about what the generator produces, and wrong assumptions there propagate silently into everything trained downstream.
That framing sets up the right answer to the design question too. If asked how you would handle a task with only a few hundred examples, a complete answer is: prompt first, bootstrap more examples from a strong model, filter hard, fine-tune something small on the result, and evaluate on the real examples you held back — never on the synthetic ones.