ModelsParticula-Classify
Labels at scale,in the hot path.
Particula-Classify assigns intent, sentiment, and categories against a taxonomy you lock in advance, at 12 ms median and 10K+ requests per second on a single node. Small enough to sit in the request path, self-hosted so the traffic never leaves.
YOUR-GPU · PARTICULA-CLASSIFY
$ POST /v1/classify · "Where is my refund?"
→ particula-classify · labels locked
→ "intent": "refund_status"
→ "sentiment": "negative" · "confidence": 0.97
✓ 200 · 12 ms · 10K+ req/s sustained
ready
PARTICULA-CLASSIFY · YOUR HARDWARE · ZERO TELEMETRY
What it doesSee it on your data
Your taxonomy, lockedLabels come from the label set you define, never free text. No drift, no novel categories appearing in production.
10K+ requests per secondSustained on a single node, with batching handled by the runtime. Built for ticket queues, chat routing, and event streams.
Confidence you can route onScores are calibrated, so a 0.97 means 0.97. Send low-confidence items to a human queue with a one-line rule.
Retrains on your labelsA few hundred corrected examples usually move the needle. The included tuning job runs overnight on the same GPU.
Spec sheetCompare all six
Context window8K tokensSized for tickets, messages, and reviews, not book-length input.
Median latency12 ms medianBatch 1 on the reference single-GPU deployment. Our own measurement: the shipped harness reproduces it on your hardware.
InterfaceOpenAI-compatible RESTServed by vLLM, deployed with Docker Compose or Helm. Drop-in for existing OpenAI client code.
Hardware1× 16 GB GPUT4 class or better. Sizing reviewed before the pilot.
Fine-tuningLoRA on your dataTuning jobs are included with Enterprise. A few hundred labeled examples from your pipeline is usually enough to move the needle.
How deployment worksStart at step 01
01Discovery call
30 minutes, no slides. Your task, your data shape, and the security constraints around it. If this model is the wrong fit, we say so on the call.02Pilot install
Containerized deployment to staging in your VPC or on-prem, on the same stack as Lumen and Notetaker. The eval harness runs on your data and the numbers go into the pilot agreement.03Production rollout
Runbook handoff, operator training, and 90 days on call. After that it is yours to operate, no vendor dependency.More modelsAll six models

Request a demo
“Bring a hundred real examples from your pipeline. We run the model on them live, the harness scores the output in front of you, and you decide with numbers instead of a sales deck.”
Sebastian Mondragon, Founder of Particula Tech
Request a demo30 minutes · no sales deck