Small models, one job each.On your hardware.
Six specialized models, each built for one narrow task. Every one runs on a single GPU inside your network, behind an OpenAI-compatible API.
Six models, one runtime.
Each trained for a single job.
Pick the model that matches the task. They share a serving stack, a deployment path, and one API, so licensing a second one adds no operational surface.
Particula-JSON
Documents and free text into schema-valid JSON. Constrained decoding: malformed output is impossible.
99.8% SCHEMA PASS02Particula-Classify
Intent, sentiment, and categories against your locked taxonomy, at single-digit milliseconds.
10K+ REQ/S03Particula-Code
Generation that compiles and runs your tests before it answers. Responses ship with the tests they passed.
PY · TS · GO · RUST04Particula-Healthcare
Diagnoses, medications, and codes from clinical notes. PHI never leaves the building.
HIPAA-READY05Particula-Legal
Key clauses and risk flags from contracts, with span-level citations. Privilege stays intact, on-prem.
PRIVILEGE-SAFE06Particula-Finance
Figures from filings, statements, and decks, down to the footnote, with a source span on every number.
99%+ EXTRACTIONRent a generalist,
or own a specialist.
On a narrow task, a small model is faster, steadier, and easier to test than a general one doing the same work. It also fits on hardware your team already knows how to buy.
One narrow job, done properly
Each model is trained on a single task and nothing else. For that task, small beats general: faster, more consistent, easier to test.
Small enough to own
Every model fits on a single commodity GPU, so the whole family runs on hardware you already know how to buy.
Your data never leaves
Inference happens inside your network: VPC, on-prem, or air-gapped. No vendor cloud in the request path, zero telemetry out.
Boring to operate
One vLLM runtime, an OpenAI-compatible API, Docker or Helm. The same deployment stack as Lumen and Notetaker, on purpose.
A flat annual license.
No per-token meter.
Licensing is per model, sized by deployment scope. Quotes are written after the demo, once the scope is clear.
We measure on your data,
not on a public leaderboard.
The eval harness ships inside every container. The numbers that decide the rollout come from your samples, on your hardware, and go into the pilot agreement in writing.
- 01
No leaderboard theater
Public leaderboards reward generality. These models are narrow on purpose, so we do not publish rankings against frontier models. The only benchmark that matters is your task.
- 02
The harness ships with the model
Every container includes the eval harness and our held-out test suites. Point it at your samples and get accuracy, latency, and throughput on your own hardware.
- 03
Thresholds in the contract
The pilot defines target metrics on your data, in writing. The rollout decision is measured against those numbers, not a demo impression.
“Bring a hundred real examples from your pipeline. We run the model on them live, the harness scores the output in front of you, and you decide with numbers instead of a sales deck.”