[home]   [coding projects]   [research projects]   [research interests]


Synthetic Data

A manual for curriculum-guided local corpus generation, filtering, and evaluation.

This project is a corpus-construction pipeline, not merely a loop that asks a language model for more text. It separates planning, curriculum construction, generation, filtering, and evaluation so that each stage can be inspected or replaced independently.

The guiding idea is that a seed example contains more than wording. It has an artifact type, a domain, skills, useful transformations, an audience, a difficulty level, stylistic constraints, and a structural pattern. The pipeline tries to extract those properties first and generate from that representation rather than repeatedly paraphrasing the original text.

Contents

1. Pipeline
2. Planning
3. Generation modes
4. Filtering
5. Running the generator
6. Evaluating the corpus

1. Pipeline

Seed examples become plans and a curriculum before any final records are generated.
Seed examples become plans and a curriculum before any final records are generated.

The generator can use Ollama, a local Hugging Face model, an arbitrary command-line process, or a mock backend for testing. Provider choice is isolated behind the same client interface, so the planning and filtering logic does not depend on one particular model runtime.


2. Planning before generation

Every seed receives a structured plan. The plan is intentionally richer than a topic label:

generator/make_synthetic_dataset.py
@dataclass
class Plan:
    source_id: str
    artifact_type: str
    domain: str
    skills: List[str]
    useful_operations: List[str]
    avoid_operations: List[str]
    possible_audiences: List[str]
    difficulty: str
    style_constraints: List[str]
    structure: List[str]
    notes_for_generator: str

If model-produced planning fails, the modular planner has a deterministic fallback with generic skills, useful operations, audiences, and a setup/development/conclusion structure. A global curriculum is then built from the collection of seed plans. That summary supplies the “frontier” jobs later in the pipeline.


3. Three generation modes

Local, recombine, and frontier jobs move progressively farther from an individual seed.
Local, recombine, and frontier jobs move progressively farther from an individual seed.

A local job chooses one seed and one operation that its plan says is useful. A recombine job samples two or three plans and asks for a new record that combines useful traits without copying their wording. A frontier job has no direct seed at all and samples from the global curriculum.

generator/jobs.py
for i in range(local_n):
    sid = rng.choice(seed_ids)
    plan = plans_by_seed.get(sid, {})
    op = rng.choice(_ops_from_plan(plan, fallback_ops))

for i in range(recombine_n):
    k = min(rng.choice([2, 2, 3]), len(seed_ids))
    sids = rng.sample(seed_ids, k=k)

for i in range(frontier_n):
    op = rng.choice(fallback_ops)

The requested ratios are normalized to the target record count and the jobs are shuffled before generation. This avoids generating the entire corpus in three large homogeneous blocks.


4. Filtering and duplicate rejection

A generated record must survive a common filter regardless of how it was produced. The filter checks length, banned model/meta phrases, fake-citation risk, repeated lines, repeated n-grams, similarity to seeds, and similarity to records that have already been accepted.

generator/filters.py
if wc < cfg.min_words:
    reasons.append("too_short")
if wc > cfg.max_words:
    reasons.append("too_long")

if _repeated_line_fraction(stripped) > cfg.max_repeated_line_fraction:
    reasons.append("repeated_lines")

if _repeated_ngram_fraction(stripped) > cfg.max_repeated_ngram_fraction:
    reasons.append("repeated_ngrams")

for seed in seed_texts:
    sim = jaccard(
        shingles,
        shingle_set(seed, cfg.shingle_n)
    )
    if sim > cfg.max_seed_jaccard:
        reasons.append(f"too_close_to_seed:{sim:.3f}")
        break

for old in existing_shingles:
    sim = jaccard(shingles, old)
    if sim > cfg.max_existing_jaccard:
        reasons.append(f"near_duplicate:{sim:.3f}")
        break

This stage is important because the goal is not simply to maximize generation count. The output should move away from the seed set without collapsing into near-duplicates, boilerplate, or invented citations.


5. Running the generator

The supplied start script creates a 10,000-record dataset from a seed file through a local Ollama server and includes the seed material in the compiled output:

shell
bash scripts/start_data.sh

The corresponding resume script uses the same configuration with --resume, so plans and accepted records already written to disk can be reused after an interrupted run.


6. Evaluating the resulting corpus

The repository also contains a small Transformer training harness and separate scripts for control, arxiv, and synthetic corpora. This turns corpus construction into an experiment rather than a one-way export: the generated data can be compared under the same model size, sequence length, optimization settings, and validation text.

shell
bash scripts/synthetic.sh

The training report includes train and validation loss, perplexity, bits per byte, train/validation gap, overfit ratio, loss improvement, and throughput. Those diagnostics are useful for asking whether the generated corpus changes generalization rather than merely whether the generator produced a large file.

PathPurpose
generator/planning.pySeed plans and curriculum summary.
generator/jobs.pyLocal / recombine / frontier scheduling.
generator/llm_clients.pyModel backend abstraction.
generator/filters.pyQuality, repetition, citation, and similarity rejection.
generator/make_synthetic_dataset.pyEnd-to-end generation and resume logic.
exp/exp_main.pyLanguage-model evaluation of resulting corpora.

These pages are selective technical manuals: enough source to expose the mechanism, not a mirror of the entire repository.
Last updated: September 14, 2026.