I built a synthetic data pipeline that uses an LLM to generate instruction-following datasets for fine-tuning smaller language models. It is useful when you need curated training data but do not have a labeling team. I run it on Oxlo.ai because flat per-request pricing keeps costs predictable even when I stuff long system prompts into every call.
What you'll need
- Python 3.10 or newer
- An Oxlo.ai API key from https://portal.oxlo.ai
- The OpenAI SDK installed with
pip install openai
Oxlo.ai is fully OpenAI SDK compatible, so the code below drops in without changes. You can explore model options and pricing at https://oxlo.ai/pricing.
Step 1: Connect and Test the Oxlo.ai Client
First, I set up the client and verify the connection. I use llama-3.3-70b as the workhorse because it follows instructions reliably and starts instantly with no cold starts on Oxlo.ai.
from openai import OpenAI
import os
client = OpenAI(
base_url="https://api.oxlo.ai/v1",
api_key=os.getenv("OXLO_API_KEY", "YOUR_OXLO_API_KEY")
)
response = client.chat.completions.create(
model="llama-3.3-70b",
messages=[{"role": "user", "content": "Say 'Connection OK'"}],
max_tokens=10,
)
print(response.choices[0].message.content)
Step 2: Define the Generator Prompt
The system prompt is the only manual labeling we do. It forces the model to emit strict JSON so downstream tools can consume the data without parsing surprises.
SYSTEM_PROMPT = """You are a synthetic data engineer. Your job is to generate high-quality instruction-following training pairs for fine-tuning language models.
Rules:
- Output strictly valid JSON with keys: instruction, input, output.
- The instruction must be a clear, standalone task.
- The input field provides context and can be an empty string.
- The output field is the ideal response.
- Vary tone and complexity. Do not wrap the JSON in markdown code fences.
- Do not include explanatory text outside the JSON object."""
Step 3: Generate a Single Training Pair
Next, I wrap the API call in a function that accepts a topic and difficulty. I keep the temperature at 0.8 to balance creativity with coherence.
import json
def generate_raw_pair(topic, difficulty="intermediate"):
user_message = f"Generate one training pair about: {topic}. Difficulty: {difficulty}."
response = client.chat.completions.create(
model="llama-3.3-70b",
messages=[
{"role": "system", "content": SYSTEM_PROMPT},
{"role": "user", "content": user_message},
],
temperature=0.8,
)
return response.choices[0].message.content.strip()
print(generate_raw_pair("Python list comprehensions"))
Step 4: Validate and Structure Output
LLMs sometimes return markdown or trailing commentary. I add a validator that strips fences, parses JSON, and retries up to three times before giving up.
def generate_valid_pair(topic, difficulty="intermediate", retries=3):
user_message = f"Generate one training pair about: {topic}. Difficulty: {difficulty}."
for attempt in range(retries):
response = client.chat.completions.create(
model="llama-3.3-70b",
messages=[
{"role": "system", "content": SYSTEM_PROMPT},
{"role": "user", "content": user_message},
],
temperature=0.8,
)
raw = response.choices[0].message.content.strip()
# Strip markdown fences if the model ignored instructions
if raw.startswith("
```json"):
raw = raw.split("```
json", 1)[1]
if raw.endswith("
```"):
raw = raw.rsplit("```
", 1)[0]
raw = raw.strip()
try:
parsed = json.loads(raw)
if all(k in parsed for k in ("instruction", "input", "output")):
return parsed
except json.JSONDecodeError:
continue
raise ValueError(f"Failed to generate valid JSON for topic: {topic}")
pair = generate_valid_pair("SQL JOIN statements")
print(json.dumps(pair, indent=2))
Step 5: Batch Generate and Export to JSONL
Finally, I loop over a list of topics, collect valid pairs, and stream them to a JSONL file. I add a small sleep between calls to stay polite, though Oxlo.ai handles concurrency well.
import time
topics = [
"Python decorators",
"React useEffect hooks",
"Docker compose networking",
"Git rebase vs merge",
"REST API pagination",
]
dataset = []
for topic in topics:
try:
pair = generate_valid_pair(topic)
dataset.append(pair)
time.sleep(0.5)
except Exception as e:
print(f"Skip {topic}: {e}")
with open("training_data.jsonl", "w", encoding="utf-8") as f:
for item in dataset:
f.write(json.dumps(item, ensure_ascii=False) + "\n")
print(f"Wrote {len(dataset)} examples to training_data.jsonl")
Run It
I usually run the script directly and inspect the first few lines. Here is the entry point I use and an example of the output.
if __name__ == "__main__":
print("=== Single test ===")
test = generate_valid_pair("Explain recursion", difficulty="beginner")
print(json.dumps(test, indent=2))
print("\n=== Batch job ===")
# Re-run the batch loop above to produce training_data.jsonl
Example output from the single test:
{
"instruction": "Explain the concept of recursion in programming to a beginner.",
"input": "",
"output": "Recursion is when a function calls itself to solve a smaller version of the same problem. Imagine a set of Russian dolls. To count how many dolls there are, you open the outer doll and count the dolls inside it, then add one for the doll you just opened. The function keeps calling itself on the smaller set until it reaches the smallest doll, which does not contain another doll. That smallest doll is the base case, and it stops the recursion."
}
Next Steps
Try swapping llama-3.3-70b for qwen-3-32b or deepseek-v3.2 to see which model produces the best domain-specific data for your use case. Once you have a few hundred examples, you can plug the JSONL file into a fine-tuning pipeline on Oxlo.ai or your own infrastructure.
Top comments (0)