Preregistered Jacobian-lens reliability research on Gemma 4: 25k prompts, frozen probes, public traces, a prospective transfer miss, and an answer-identity confound.
-
Updated
Jul 19, 2026 - Python
Preregistered Jacobian-lens reliability research on Gemma 4: 25k prompts, frozen probes, public traces, a prospective transfer miss, and an answer-identity confound.
Turn Chaos Into Structure. A Type-Safe AI Agent that extracts valid JSON from unstructured data using PydanticAI, FastHTML, and Gemini 2.5.
Interactive Phoenix LiveView demonstrations of the Crucible Framework - showcasing ensemble voting, request hedging, statistical analysis, and more with mock LLMs
Code for the ICML 2026 Main Track Paper "Cognitive Fatigue in Autoregressive Transformers: Formalization and Measurement"
Three small LLMs, one CPU, no cloud. Rigorous benchmarking, structured output validation, and head-to-head quality scoring via Ollama + FastAPI.
Reference implementation of CAAF — three-pillar agent framework with monotonic convergence.
TypeScript eval harness for measuring whether Grok answers stay grounded in source evidence
A Python to Prolog pipeline that bounds Elasticsearch LLM errors by grounding ES diagnostics analysis in verified metrics.
Extracts a driver profile from a call transcript behind a validating schema gate, then screens a load board before ranking by effective rate per mile.
Profile README. AI engineer working on agentic LLM systems, deterministic guardrails, and how models fail in long interactions.
Collection of LLM failure modes used on failmodes.com
Public artifact bundle for the preprint 'Lightweight Evaluation and Operational Scorecards for Tool-Using AI Agents'
Companion project for the TechnologyDig Academy tutorial on building reliable generative programs with Mellea.
Catches safety constraints and task invariants that an LLM agent's context compaction silently drops, then repairs them. Zero dependencies.
Reliability and hallucination mitigation research for tool-augmented legal AI agents using QC-Sentinel verification architecture.
Official implementation of Cascading Hallucination in Agentic RAG: The CHARM Framework for Detection and Mitigation.
CrucibleFramework: A scientific platform for LLM reliability research on the BEAM
When the agent breaks, which layer dropped the ball? Operator-facing failure-attribution taxonomy for AI agent estates: ten-way dictionary, MAST + AgentRx crosswalks, postmortem template. Part of the Spine catalog.
Map where your bolted-on AI feature breaks before customers do. A free Claude Code tool: fragility map, reliability score across six dimensions, ranked gaps, and a 30-day plan. Built by a threat-intel practitioner.
Reliability and audit-evidence testing for LLM agents - wrap any agent, assert behavior, measure determinism, check grounding, emit an audit-grade report.
Add a description, image, and links to the llm-reliability topic page so that developers can more easily learn about it.
To associate your repository with the llm-reliability topic, visit your repo's landing page and select "manage topics."