GENSCRIPT: Inference-Only Synthetic Data Generation Without Model Training
A new method, GENSCRIPT, enables synthetic data generation for tabular, temporal, and relational data without any model training. By using a statistical profile, language models, and a coding agent, it produces efficient, auditable row samplers supporting multiple data modalities.
What Changed
GENSCRIPT introduces an inference-only approach to synthetic data generation. Unlike the typical approach that requires training a generative model for each dataset and data modality, GENSCRIPT builds generators without any training by analyzing the dataset’s statistical profile and using a language model for field semantics and constraints.

Capabilities and Workflow
- No training is needed: Builds on statistical analysis and language model inference.
- Unified support for single-table, time series, and relational databases—no task-specific feature engineering required.
- Deterministically computes column types, ranges, missingness, categories, and correlations.
- Language model infers field semantics and inter-column integrity constraints.
- A coding agent compiles the profile and constraints into an executable, auditable data sampler.
Performance Benchmarks
- Builds synthetic data generators in about 2 minutes on four single-table benchmarks.
- Samples 50,000 rows in 6 seconds.
- Maintains marginal fidelity within a few points of leading training-based methods.
