You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
A high-performance request router for vLLM — smart load balancing, response caching, prefill/decode disaggregation, semantic routing, Anthropic/OpenAI API translation, and operational tooling out of the box.
🦀 A barebones inferencing server in Rust, using Hyper for HTTP logic and ORT for inferencing models. Containerized and easily hostable. Complete control over type of model and inferencing behavior!