Projects with this topic
-
LLM hallucination detector: highlights the words an AI answer was unsure about, from OpenAI-compatible token logprobs, and shows what it almost said instead. Terminal, HTML or Markdown report, CI gate that fails on low confidence. Rust CLI and library, runs locally.
Updated -
GitLab CI/CD components for LLM apps: fail the pipeline on likely hallucinations, prompts over their token budget, padded prompts, an LLM bill heading over budget, or slow and flaky models. Tested on every push.
Updated -
Local dashboard that streams a Claude, GPT or Ollama reply chunk by chunk, shows where the model paused, flags claims worth checking and prices the prompt before you send it. Works fully offline with Ollama.
Updated -
Async Rust library and CLI that uses a second LLM call to check the first: flags unsupported or false claims in text and critiques Rust code.
Updated -
Agent supervision, memory distillation, and multi-model orchestration. Kybernetes steers AI agents — catching hallucination and drift before they compound.
Updated -
A deterministic verification layer for AI systems. QWED verifies AI outputs using mathematics, symbolic reasoning, and formal methods (Z3, SMT, SymPy), creating an auditable trust boundary for agentic AI. Not generation. Verification.
Updated