Projects with this topic
-
CPU/System/Algorithm benchmarking framework
Updated -
Droid Tune-Up — an unofficial open-source evaluation harness for Factory Droid. Drives droid exec headless, hides tests and the solution from the agent, and grades only the committed worktree with deterministic behavioral tests. Not affiliated with Factory.
Updated -
Evaluation harness for measuring how well AI models perform on SysML v2 modeling tasks.
Updated -
Helm Charts for various benchmarks
Updated -
The repository contains the code used in an extensive benchmark of co-occurence based inference methods to recover the interaction structure of microbial communities from metabarcoding data (16S rDNA-seq data)
Updated -
Raw benchmark results and statistical analysis.
Updated -
Agent-shape testing harness that measures how an LLM-driven agent uses a tool's CLI, scored by an LLM judge.
Updated -
-
Scaling and complexity benchmarks for Univec.
Updated -
Benchmark suite for measuring Univec performance.
Updated -
Benchmark of KV stores available in Go. https://go-benchmark-kvstore.gitlab.io
Updated -
System utility designed to stress and monitor various hardware components
Updated -
Comparison of warmup times for different runtimes in AWS Lambda.
Updated -
This is a demonstration of the Whetstone Benchmark.
Updated -
Something the world really doesn't need.
Updated -
-
A simple and configurable UI to monitor the timeframe (ms) and the framerate (fps) in Unity3D.
Updated -
Measure performance of various languages in calculating cosine similarity of vectors (README will be updated...)
Updated -
Research study regarding the effect of docker image layers on performance.
Updated