Models
Small models, trained for one job.
Self-hosted on your hardware.
Six specialized models, each built for a single narrow task: extraction, classification, code, and three regulated domains. Every one runs on a single GPU inside your network, behind an OpenAI-compatible API, with no vendor cloud and no per-token meter. Deployed on the same stack as Lumen and Notetaker.
Six modelsRequest a demo
01Particula-JSON
Documents and free text into schema-valid JSON. Constrained decoding: malformed output is impossible.99.8% SCHEMA PASS02Particula-Classify
Intent, sentiment, and categories against your locked taxonomy, at single-digit milliseconds.10K+ REQ/S03Particula-Code
Generation that compiles and runs your tests before it answers. Responses ship with the tests they passed.PY · TS · GO · RUST04Particula-Healthcare
Diagnoses, medications, and codes from clinical notes. PHI never leaves the building.HIPAA-READY05Particula-Legal
Key clauses and risk flags from contracts, with span-level citations. Privilege stays intact, on-prem.PRIVILEGE-SAFE06Particula-Finance
Figures from filings, statements, and decks, down to the footnote, with a source span on every number.99%+ EXTRACTIONWhy small modelsNeed something custom?
A frontier model is a generalist you rent.
These are specialists you own.
These are specialists you own.
One narrow job, done properlyEach model is trained on a single task and nothing else. For that task, small beats general: faster, more consistent, easier to test.
Small enough to ownEvery model fits on a single commodity GPU, so the whole family runs on hardware you already know how to buy.
Your data never leavesInference happens inside your network: VPC, on-prem, or air-gapped. No vendor cloud in the request path, zero telemetry out.
Boring to operateOne vLLM runtime, an OpenAI-compatible API, Docker or Helm. The same deployment stack as Lumen and Notetaker, on purpose.
SpecificationsReproduce these at the demo
MODELTASKCONTEXTLATENCYMIN GPU
Particula-JSONExtraction to a locked JSON schemaCONTEXT32K
LATENCY41 ms
MIN GPU1× 24 GB
CONTEXT8K
LATENCY12 ms
MIN GPU1× 16 GB
CONTEXT64K
LATENCY380 ms
MIN GPU1× 48 GB
CONTEXT32K
LATENCY55 ms
MIN GPU1× 48 GB
CONTEXT64K
LATENCY62 ms
MIN GPU1× 48 GB
CONTEXT32K
LATENCY48 ms
MIN GPU1× 24 GB
Latency is the median at batch 1 on our reference single-GPU deployment. These are our own measurements, not third-party results. The same harness ships in every container, so you can reproduce every number on your hardware and your data.
How we benchmarkBring your samples
01No leaderboard theaterPublic leaderboards reward generality. These models are narrow on purpose, so we do not publish rankings against frontier models. The only benchmark that matters is your task.
02The harness ships with the modelEvery container includes the eval harness and our held-out test suites. Point it at your samples and get accuracy, latency, and throughput on your own hardware.
03Thresholds in the contractThe pilot defines target metrics on your data, in writing. The rollout decision is measured against those numbers, not a demo impression.
Pricing
A flat annual license per model.
The per-token meter is gone, because the tokens are yours.
The per-token meter is gone, because the tokens are yours.
Request a quoteWritten quote within 48 hours of the demo · valid 90 days
$0per token, per seat, per queryInference runs on your GPU, so usage costs nothing extra. Volume changes your hardware bill, not your license.
1flat annual license per modelSized by deployment scope, not usage. Includes updates, security patches, and the eval harness for the license period.
6models on one runtimeModels share the same serving stack, so licensing several earns a bundle discount and adds no operational surface.

Request a demo
“Bring a hundred real examples from your pipeline. We run the model on them live, the harness scores the output in front of you, and you decide with numbers instead of a sales deck.”
Sebastian Mondragon, Founder of Particula Tech
Request a demo30 minutes · no sales deck