Projects with this topic
Sort by:
-
Put a queue, deduplication, a circuit breaker and a dead-letter queue in front of any LLM API (Anthropic, OpenAI, llama.cpp, vLLM), so repeated prompts cost one call and outages fail fast. Rust library, HTTP server and CLI.
Updated -
Curated list of LLM infrastructure tools that hold up in production: failure handling, observability, cost control and serving.
Updated