Projects with this topic
-
Zambo: the cross-AI execution layer. 136 native MCP tools over one endpoint, with a verifiable execution receipt for every call (AER-1 open draft). Free: generous free tier, no account. Day Pass $0.99 USDC on Base for 24h; Zambo Pass $29/month ($19/month for verified $ZAMBO holders). https://zambo.dev
Updated -
Collection of Verification Tasks
Updated -
Open dataset: bitHuman avatar render speed (× real time) by device, from docs.bithuman.ai/performance. Code Apache-2.0, data CC BY 4.0.
Updated -
Droid Tune-Up — an unofficial open-source evaluation harness for Factory Droid. Drives droid exec headless, hides tests and the solution from the agent, and grades only the committed worktree with deterministic behavioral tests. Not affiliated with Factory.
Updated -
CPU/System/Algorithm benchmarking framework
Updated -
-
Raw benchmark results and statistical analysis.
Updated -
GitLab CI/CD components for LLM apps: fail the pipeline on likely hallucinations, prompts over their token budget, padded prompts, an LLM bill heading over budget, or slow and flaky models. Tested on every push.
Updated -
Объектив · Исследование распознавания техники: эталон, прогоны пайплайнов, метрики, лидерборд, выводы
Updated -
Evaluation harness for measuring how well AI models perform on SysML v2 modeling tasks.
Updated -
Model Verification Layer is a modular system for evaluating, comparing, and validating the behavior of large language models using structured benchmarks, logic consistency checks, cross-model consensus analysis, and policy-aware constraints. It provides a transparent framework for understanding how different AI models perform under identical conditions, enabling more reliable model selection, safer deployment, and user-adaptive decision making. https://roxanneardary.com/model-verification-layer/
Updated -
Public evidence datasets, claim atlases, and reproducible benchmarks for ClaimBound.
Updated -
-
An Extensible Benchmark Framework for Real-Time Applications
Documentation: https://rt-bench.gitlab.io/rt-bench/
Updated -
-
-
Helm Charts for various benchmarks
Updated -
The repository contains the code used in an extensive benchmark of co-occurence based inference methods to recover the interaction structure of microbial communities from metabarcoding data (16S rDNA-seq data)
Updated -
Agent-shape testing harness that measures how an LLM-driven agent uses a tool's CLI, scored by an LLM judge.
Updated -
Empirical validation of C4 geometric defense against 16 Agents of Chaos. 550 adversarial prompts. 4 defense systems. 96.7% block rate. LLM validation on GPT-4o-mini + Mistral 7B. MIT.
Updated