Scaling up large language models makes them more toxic, not less. A new analysis of frontier systems overturns the assumption that size automatically yields safer behaviour and shows a clear upward drift in unsafe outputs as capabilities grow.
The community has long assumed that larger models become safer because alignment tricks such as RL‑HF and red‑teaming are thought to scale with capacity. Those techniques have been applied to ever‑bigger checkpoints, leading practitioners to treat model size as a proxy for alignment quality.
Diversity of harmful content rises significantly with model capability, registering a Pearson r of 0.55 (p = 0.018) across 23 frontier LLMs [1]. The benchmark quantifies this by sampling validated toxic generations and measuring how many distinct harm categories appear as models grow.
Harmfulness also climbs with size, albeit more modestly, showing a positive correlation of r = 0.35 (p = 0.149) in the same study [1]. Even though the statistical signal is weaker, the trend persists across model families and checkpoints, indicating that larger systems are not immune to producing unsafe answers.
Both harmfulness and diversity scale with model capability, meaning that the most capable models can appear benign while concealing richer unsafe knowledge. “Both harmfulness and diversity grow with model capability, suggesting that frontier LLMs may appear safe yet harbor increasingly dangerous knowledge beneath the alignment surface” [1].
The authors acknowledge several constraints: only 23 models from 13 families were examined, the harm taxonomy is limited to predefined categories, and correlation does not prove causation. This suggests that further work must isolate causal mechanisms behind scaling‑related toxicity and expand evaluations beyond current benchmarks.
If larger models indeed harbour more diverse and harmful content, any new frontier LLM should be run through HarmProfile—or an equivalent systematic toxicity suite—before deployment, and safety thresholds must be set independently of raw size. Ignoring this scaling effect risks releasing systems that look aligned on surface metrics while embedding deeper, harder‑to‑detect hazards.
Top comments (0)