Projects with this topic
-
10 hands-on demos for agentic AI security: blind verification, AIBOM, eval invariants, authority confinement, recon/malware/LFI/SSRF/scan — offline only.
Updated -
Single-header C++17: detect and scrub PII (email, phone, SSN, card numbers, API keys) and score prompt-injection risk. Part of llm-cpp.
Updated -
Integrity monitoring for the files your AI coding agent obeys — detects drift, invisible characters, bidi controls, and homoglyphs in CLAUDE.md, .cursorrules, settings.json and friends.
Updated -
Field guide for LLM prompt injection: detection, categories, evasion, and defense mapping.
Updated -
Scan text and files for hidden steganography and prompt injection before they reach an LLM. Clean what you can, quarantine the rest.
Updated -
Evaluate AI web-browsing agents against adversarial pages. Franklin et al. (2026) attack-class taxonomy plus StegOFF blocking.
Updated -
A skill for AI coding agents that scaffolds a safe, multi-model chatbot for Telegram or Discord. Supports Claude, GPT, Gemini, and OpenAI-compatible backends. Nine safety layers on by default. Named for R. Daneel Olivaw from Asimov.
Updated -
For testing or trolling LLM based pentesting frameworks.
Updated