Projects with this topic
-
Python toolkit for benchmarking RAG quality on Salesforce Data Cloud. Includes retrieval metrics (Hit Rate, MRR, NDCG), generation evaluation, and optional Claude Code skills.
Updated -
ValidationOS is an open-source, local-first AI assurance platform that transforms AI performance into proven results. It provides tools for validating, optimizing, certifying, and monitoring AI models across edge hardware environments through automated testing, compatibility analysis, performance benchmarking, observability, and compliance workflows.
Updated -
Automated LLM Benchmarking on GPU - tokens/sec, latency percentiles, VRAM profiling, multi-format support (HuggingFace, GGUF, GPTQ)
Updated -
A minimalistic and portable unit testing and benchmarking framework for C.
Updated -
PTA 2026. Task Benchmarks
Updated -
-
Benchmarks data processing tools
Updated -
MultiNativQA is Multilingual Native question-answering (QA) dataset consisting of 64k QA pairs in seven extremely low to high resource languages, covering 18 different topics from nine different regions. Paper: https://arxiv.org/pdf/2407.09823. Project: https://nativqa.gitlab.io
Updated -
Cristian Vasu Data Portfolio / Database Performance Benchmarking with YCSB - Cassandra vs PostgreSQL
Benchmarked columnar vs row-oriented databases using the Yahoo Cloud Server Benchmark (YCSB). Compared throughput, latency, and error rates under different workloads. Analyzed trade-offs between NoSQL (Cassandra) and relational (PostgreSQL) databases with supporting graphs and configuration files.
Updated -
This C++ sample project uses meson to manage builds and already has auto-formatting, testing, linting, benchmarking, code coverage, static analysis and documentation generation.
Updated -
-
Benchmarking framework for machine learning with fNIRS
Updated -
QAPerf is a tool for evaluating the performance of quality assessment (QA) methods. It automates the generation of standard reports, including correlation tables, leave-one-group-out correlation boxplots, and other benchmarking visualizations. Designed for researchers and practitioners, QAPerf streamlines the analysis of QA metrics, ensuring reliable and reproducible evaluations.
Updated -
Command-line benchmarking tool. Runs commands, times them, ranks them. Also reports memory used, context switches, and other metrics.
Updated -
Server fibonacci benchmarks in three different technologies: node-js, vapor, php
Updated -
A Live Evaluation of Computational Methods for Metagenome Investigation
Updated -
SecurityPerf is a tool designed for benchmarking production workloads. In doing so, it makes measuring the impact of security programs on production workloads easy.
Updated -
VTmark results, Ansible roles and other scripts.
Updated -
Pew Pew is a simple HTTPS benchmarking tool written in Go.
Updated