Synthetic Data Generation

Post-training runs on data that often does not exist: instructions nobody wrote, rare cases nobody recorded, records nobody may share. Generating that data — with a model, a simulator or a statistical recipe — is now a standard step, and it has failure modes of its own.

Large Language Models: From Transformers to Frontier Models

Everything in this module runs on data. Fine-tuning needs examples of the behaviour you want; alignment needs pairs of answers to compare. And here is the awkward truth nobody mentions in the papers: for most real projects that data does not exist, or cannot be shared, or would take a year and a budget you do not have to collect. Synthetic data — examples produced rather than observed — is how teams get past that. It has quietly gone from a last resort to a standard step, and it brings failure modes of its own that are worth knowing before you rely on it.

What synthetic data is for

Four constraints push teams toward generating data, and they are worth separating because they call for different techniques.

Coverage of rare events. The cases that matter most are often the ones you have least of: the fraudulent transaction, the near-collision, the equipment failure. A model trained on what happens usually will be weakest exactly where the stakes are highest, and you cannot ethically arrange more crashes to fix that.

Privacy and access. Health, financial and legal records are the data you most want and least may use. A generated dataset with the same statistical shape as the real one can be shared across teams, used in testing, or handed to a vendor, without exposing anyone — provided it is genuinely generated and not quietly copied, which is a claim that has to be checked rather than assumed.

Imbalance. When one class is a fraction of a percent of the data, a model can score well by ignoring it entirely. Adding plausible examples of the minority class changes what the model is rewarded for.

Speed. Real data arrives on its own schedule — collection, labelling, review. Synthetic data arrives when you ask for it, which is what makes it so useful early, when you are still deciding whether a model is worth building at all.

In post-training, the generator is a model

Inside an LLM pipeline, synthetic data usually means a strong model producing training data for a weaker one. You met the logic in the pruning and distillation lesson: a very large model is expensive to serve but excellent as a source of examples, and the examples outlive it.

The pattern that established this starts from a seed set of a few hundred hand-written tasks, each with an instruction and one worked example. A base model is shown a few of them and asked to invent another instruction of the same kind; given an accepted instruction, it is asked to produce the input and the response. Filtering removes duplicates, malformed items and tasks that are too close to ones already in the pool. Iterate, and a couple of hundred human-written seeds become a large instruction-tuning set — which is then used to fine-tune a model that could not have followed those instructions beforehand.

Two details make this work rather than collapse into noise:

  • Few-shot prompting is enough. A base model cannot reliably follow a zero-shot instruction — that is what post-training is for — but it imitates a handful of examples perfectly well. Generation does not require an already-aligned model.
  • Filtering is not optional. It is the step that turns generation into a dataset. Deduplication, format validation, and wherever possible an automatic correctness check, decide the quality of everything downstream.

This is now the normal way to build a narrow capability. To put a natural-language interface on an existing system, you describe what the system can do, have a strong model write many plausible user phrasings and correct responses, filter, and fine-tune a small model on the result. Nobody writes the thousands of examples by hand.

A small set of seed examples feeding a generator model, whose output passes through filtering before joining a training set used to fine-tune a smaller model, with the pool feeding back into generation
The generation loop: a few hand-written seeds, a strong model that proposes more, a filter that decides what survives, and a growing pool that feeds the next round.
Why a model can teach a smaller model anything at all

It looks like something for nothing — new training signal from a model that has seen no new data. It is not. An over-parameterised model has absorbed patterns from an enormous corpus that a small model trained directly would never extract. Generation makes that latent knowledge explicit, in a form a small model can learn from efficiently. Nothing new enters the system; it is rearranged into a more learnable shape. Which is also why the ceiling is the teacher: a student trained this way inherits the teacher's blind spots along with its strengths.

Outside language, the toolkit is statistical

Not all synthetic data comes from a generative model, and an AI engineer is expected to know the cheaper options.

Sampling from specified distributions is the most direct. You write down how you believe each field behaves — an age roughly normal around some mean, a salary that rises with age plus noise, a category drawn from a fixed list — and draw as many records as you like. A seasonal series is a sine wave plus Gaussian noise. The strength is total control and total transparency: you know exactly what relationships exist, because you put them there. The weakness is the same sentence read differently — the data contains no structure you did not think of, so it is excellent for tests, prototypes and load simulation, and poor as a stand-in for reality.

Interpolation between real examples handles imbalance. The best-known method, SMOTE, takes a minority-class point, finds its nearest neighbours among other minority points, and creates a new point somewhere on the line between them. Rather than duplicating rare rows — which teaches a model nothing new — it fills in the space around them, giving the classifier a smoother boundary to find. It assumes that the straight line between two real examples contains plausible examples, which holds for continuous numeric features and fails for categorical ones, where the midpoint of two categories is meaningless. It also degrades when the minority class is scattered or overlaps the majority, since interpolating across a gap plants new points in territory that belongs to the other class.

Learned generative models handle the cases the first two cannot. For tables, a conditional GAN learns the joint distribution of mixed numeric and categorical columns and generates whole rows, conditioning on discrete fields so it respects dependencies a simpler method would violate — that a ten-year-old is not a department head, for instance. It is slower, needs tuning, and is opaque, but it captures non-linear multi-column structure that neither hand-specified distributions nor interpolation can.

For images and audio, augmentation comes first. Rotating, cropping, flipping and colour-shifting an image, or pitch-shifting and adding room noise to a recording, produces new training examples whose labels you already know. The content is unchanged; what changes is the range of conditions under which the model must still recognise it. Full generation — synthesising entirely new images with a generative model — is the heavier option, used where even augmented real data is too scarce.

Specified distributionsInterpolation (SMOTE)Learned generator
How it worksDraw from distributions you write downInterpolate between real neighboursLearn the joint distribution, then sample
Captures dependenciesOnly the ones you encodeLocal and linear onlyNon-linear, across many fields
Mixed data typesManual encodingNumeric featuresHandled directly
CostNegligibleLowHigh, needs tuning
PrivacyNo real data involvedBuilt from real pointsCan memorise — must be tested
Good forTests, prototypes, simulationClass imbalanceRealistic, shareable datasets

Judging whether it worked

Generated data is not automatically good data, and a single score will not tell you. Check three things, because a dataset can pass any one of them and fail badly on another.

Realism — does it look like the real thing? Compare distributions field by field: a Kolmogorov–Smirnov test for numeric columns, a chi-squared test for categorical counts. These are cheap and catch gross distortions, but they only examine one field at a time, so passing them means the margins match, not that the relationships do.

Utility — does it work? Train a model on the synthetic data and evaluate it on held-out real data. If it comes close to a model trained on real data, the synthetic set carries the structure that matters. This is the test that counts, because it measures the thing you actually want. A projection of both datasets into two dimensions makes a quick visual check: the two clouds should overlap.

Privacy — did it copy? Search each synthetic record for its nearest real neighbour. Records that sit implausibly close to a real one suggest memorisation rather than generation, which defeats the purpose entirely in a regulated domain.

Generated does not mean anonymous

The assumption that synthetic data is automatically safe to publish is both common and wrong. A generative model trained on sensitive records can overfit and reproduce individual examples nearly verbatim, and a rare record is the most likely to be memorised — exactly the record whose exposure matters most. Whether a dataset is safe to release is an empirical question, answered by nearest-neighbour distance checks and membership-inference tests, not by the fact that a model produced it.

Where it goes wrong

Mode collapse. A generative model can learn to produce a narrow slice of the distribution and repeat it — a few common patterns, with the tails missing. The irony is sharp: you often generate data because the rare cases are underrepresented, and a collapsed generator will faithfully leave them out. Small, imbalanced source datasets make this more likely.

Memorisation, as above: the privacy failure that looks like success until someone measures it.

Implausible samples. Interpolation methods have no notion of what is possible, only of what is between two points. In a high-dimensional or noisy space the midpoint of two valid records can be a record that could not exist.

Bias amplification. This is the subtlest and the most damaging. A generator reproduces the patterns in its training data, including the ones you would not have chosen — an underrepresented group, a historical labelling bias, a skew in who appears in which role. It can also sharpen them, since a model trained to produce likely examples will produce the majority pattern more reliably than the data did. Correcting the training data is not enough; the generation process itself has to be audited.

Drift toward the generator. When synthetic data is used at scale to train models whose outputs are then used to generate more data, successive generations can narrow toward what the generator finds easy, losing the variety of the original distribution. Keeping real data in the mix is the standard defence.

None of this makes synthetic data a bad idea — the alternative is usually no data at all. It makes it a tool with preconditions: generate deliberately, filter hard, test on real data, and check what the generator memorised before anyone else sees it.

MediumSynthetic data

Why is filtering, not generation, the step that determines synthetic data quality?

MediumSynthetic dataevaluation

A synthetic dataset passes distribution tests on every field. Why isn't that sufficient?

HardSynthetic datafailure modes

Why is mode collapse particularly damaging for the reason people usually generate data?

MediumSynthetic dataprivacy

Why can't you assume a generated dataset is safe to share?