Notes from production AI.Not marketing content.
What we learn shipping AI onto client hardware: self-hosting economics, model choices, and security.
Roughly 45 to 48 percent of agent failures close with a confident completion claim. LLM judges catch almost none of them. Here is what actually does.

Everything we have
written down.
Every post comes out of work we did: a client system, an internal build, or a benchmark we ran. Filter by topic above, or read straight through.

AI Agent Says Done But Did Nothing: Detect False Success
Roughly 45 to 48 percent of agent failures close with a confident completion claim. LLM judges catch almost none of them. Here is what actually does.
AUG 13, 2026
OpenTelemetry GenAI Semantic Conventions: 0 of 63 Are Stable
All 63 gen_ai attribute keys still read stability: development. Here is the deprecation map, the package flip that blanks dashboards, and what to pin.
AUG 13, 2026
Presidio vs GLiNER for PII Redaction Before an LLM in 2026
Presidio left Microsoft in June 2026, and only 3 of its 81 entity types come from a model. Which PII classes each detector misses, and where the hop belongs.
AUG 13, 2026
AI Liability Insurance Exclusions: What CG 40 47 Removes
Endorsement CG 40 47 01 26 has been attachable at general liability renewals since January 2026 and strips AI claims from both coverage parts. What to check.
AUG 12, 2026
Vector Database Security: Qdrant, Milvus, Weaviate, Chroma
Qdrant 1.19.0 treats an empty api_key as unset. Milvus publishes its object store on 9000 with default credentials. Four CVEs set the version floor.
AUG 12, 2026
RAG Returns Outdated Documents: Effective Dates and Ranking
Naive RAG answers 58% of version-sensitive questions correctly; a version-aware pipeline hits 90%. Here is the effective-dating schema that closes the gap.
AUG 12, 2026
AI Due Diligence Data Room Review: From Index to Request List
Preregistered testing put hallucination at 17 to 33 percent in retrieval-backed legal AI. Here is the four-pass data room review that survives a dispute.
AUG 11, 2026
AI for AML Alert Triage: Cutting a 95% False Positive Rate
Rule-based AML monitoring runs 90 to 95% false positives at 30 minutes an alert. Here is the triage pipeline, the evidence contract, and the validation path.
AUG 11, 2026
How to Automate Contract Review With AI: A 5-Stage Pipeline
Published benchmarks put clause-level risk detection at 0.644 F1, not the 95% on product pages. The five-stage pipeline, and the exact point it must stop.
AUG 11, 2026Not ready to talk?
What we learn shipping AI systems, once a week. No announcements, no fluff.