Projects with this topic
-
Collection of Verification Tasks
Updated -
Kevlar Benchmark: OWASP Top 10 for Agentic Apps (AI-Agents) 2026 a Red Team Benchmark.
Updated -
Public evidence datasets, claim atlases, and reproducible benchmarks for ClaimBound.
Updated -
Evaluation harness for measuring how well AI models perform on SysML v2 modeling tasks.
Updated -
Agent-shape testing harness that measures how an LLM-driven agent uses a tool's CLI, scored by an LLM judge.
Updated -
Raw benchmark results and statistical analysis.
Updated -
Model Verification Layer is a modular system for evaluating, comparing, and validating the behavior of large language models using structured benchmarks, logic consistency checks, cross-model consensus analysis, and policy-aware constraints. It provides a transparent framework for understanding how different AI models perform under identical conditions, enabling more reliable model selection, safer deployment, and user-adaptive decision making. https://roxanneardary.com/model-verification-layer/
Updated -
An Extensible Benchmark Framework for Real-Time Applications
Documentation: https://rt-bench.gitlab.io/rt-bench/
Updated -
A diagnostic framework for measuring LLM vulnerability to Affective Contextual Erosion (ACE) and related liminal attack vectors. Delirium is not an exploitation tool. It is a standardized benchmark designed to detect the precise moment when a language model's attention weights shift from serving a system prompt to serving an emergent interpersonal pattern — before harm occurs.
Updated -
CPU/System/Algorithm benchmarking framework
Updated -
A benchmark library with statistical analysis and plotting capabilities in C++. https://cppstatbench.musicscience37.com/
Updated -
Red Team AI Benchmark: Evaluating LLMs for authorized offensive-security tasks. Red Team AI Benchmark is a CLI model-evaluation benchmark. It measures how LLMs understand and respond to red-team questions and security scenarios; it is not a tool for carrying out those activities. Version 2 uses a rubric-based dataset instead of judging answers only against one golden response.
Updated -
Mini benchmarking suite — performance testing utilities.
Updated -
Empirical validation of C4 geometric defense against 16 Agents of Chaos. 550 adversarial prompts. 4 defense systems. 96.7% block rate. LLM validation on GPT-4o-mini + Mistral 7B. MIT.
Updated -
The repository contains the code used in an extensive benchmark of co-occurence based inference methods to recover the interaction structure of microbial communities from metabarcoding data (16S rDNA-seq data)
Updated -
Automated LLM Benchmarking on GPU - tokens/sec, latency percentiles, VRAM profiling, multi-format support (HuggingFace, GGUF, GPTQ)
Updated -
-
Scaling and complexity benchmarks for Univec.
Updated -
Benchmark suite for measuring Univec performance.
Updated -
A Java console application that solves a sorting problem
Updated