SLM-125M — a legal language model trained from scratch

A 125.8M-parameter Llama-style base model pretrained from scratch on 2.04B tokens of US case law, SEC filings, and educational web text (val perplexity 11.35). It continues text — it is not a chatbot and has no factual grounding, so treat every completion as stylistic, not authoritative.

0 1.5
10 256
0 1
0 100
Examples