2 902 952 libros electrónicos en 111 idiomas
¿No le conviene? No hay problema. Puedes devolver los artículos hasta 30 días
No se equivocará con un vale de regalo. El destinatario puede elegir cualquier producto de nuestra oferta.
Hasta 30 días para devoluciones
Master LLM Inference and Scale Your AI Infrastructure
In 2026, inference spend surpassed training spend across the tech industry. The engineers who can maximize tokens per second on H100, H200, and B200 GPU fleets are the most valuable specialists in AI. Inference at Full Throttle turns complex GPU performance engineering into a reproducible, highly practical discipline.
Written by the ChatVariety Team-an elite collective of ML infrastructure engineers and vLLM contributors-this book provides the exact mathematical formulas and production configurations needed to run large language models at extreme scale without breaking the bank.
What You Will Master:Stop wasting millions on sub-optimal cloud GPU allocations. Learn how to design, benchmark, and operate multi-tenant, high-throughput, and ultra-low-latency LLM serving architectures today.
¡Hola! Soy Libroamiko, tu asesor de libros.
¿Cómo puedo ayudarte?