Synthetic Data Generation

Parent: data analysis · researched 2026-05-30T22:57:15.891Z· 23 sources · 9 concepts · skill da-35-synthetic-data-generation

Synthetic data is artificial data produced by a model fit to real data, designed to reproduce the real data's statistical properties (marginals, correlations, joint structure) without being a copy of

Synthetic Data Generation

When to use this skill

1. Why synthetic data

2. Tabular synthesis methods

3. Deep generative methods

4. Class imbalance: resampling vs generative

5. Differentially private synthesis

6. Tools & frameworks

7. Evaluation: fidelity vs utility vs privacy

8. Text & image synthesis (overview)

9. Regulatory context

Methodology (end-to-end pipeline)

Practical patterns

Anti-patterns

Troubleshooting

References

Children

Frontier under this node: Class imbalance: SMOTE/ADASYN vs generative, Deep generative methods (GANs, VAEs, diffusion/TabDDPM), Differentially private synthesis (DP-GAN, PATE-GAN, PrivBayes, MST, SmartNoise), Fidelity vs utility vs privacy evaluation (TSTR, DCR, membership inference), Regulatory context (GDPR, ICO, NIST), Tabular synthesis methods (Gaussian copula, CTGAN, TVAE, CopulaGAN, CART/sequential), Text and image synthesis overview, Tools and frameworks (SDV, SDMetrics, synthcity, synthpop), Why synthetic data (privacy, augmentation, testing, rebalancing)

← the whole tree · 3D view· how to read this page