LIBRISTO
LIBROAMANTO
obligatorio
Entre a formar parte de una comunidad de amantes de los libros del mundo entero y acceda a un sinfín de ventajas. Crear una cuenta gratis
0
Envío gratuito con Zásilkovna para compras superiores a 59.99 €
Mensajería SEUR 4.99 Mensajería GLS 7.99 Mensajería Correos 5.49 Mensajería DHL 5.49 Punto SEUR 3.99

Envío gratis a partir de 69,99 euros.

Inference at Full Throttle

LLM serving performance with vLLM, quantization, KV cache tuning and speculative decoding

Idioma InglésInglés
Libro Tapa blanda
Libro Inference at Full Throttle ChatVariety Team
Código Libristo: 53520833
Editores Independently published, agosto 2026
Master LLM Inference and Scale Your AI InfrastructureIn 2026, inference spend surpassed training spe... Descripción completa
? points 28 b Próximamente Próximamente Nuevo Nuevo
11.59
Reaprovisionamiento previsto Lanzamiento 17. 08. 2026

Hasta 30 días para devoluciones

Master LLM Inference and Scale Your AI Infrastructure

In 2026, inference spend surpassed training spend across the tech industry. The engineers who can maximize tokens per second on H100, H200, and B200 GPU fleets are the most valuable specialists in AI. Inference at Full Throttle turns complex GPU performance engineering into a reproducible, highly practical discipline.

Written by the ChatVariety Team-an elite collective of ML infrastructure engineers and vLLM contributors-this book provides the exact mathematical formulas and production configurations needed to run large language models at extreme scale without breaking the bank.

What You Will Master:
  • GPU Memory Mathematics: Derive the exact KV cache formula from first principles to budget HBM memory perfectly.
  • Advanced Quantization: Deploy FP8, AWQ INT4, GPTQ, and MXFP4 based on real-world throughput and quality trade-offs.
  • Serving Stack Optimization: Fine-tune vLLM, SGLang, TensorRT-LLM, and TGI for enterprise workloads.
  • Ultra-Fast Decoding: Implement speculative decoding, EAGLE-class self-speculation, and KV cache prefix caching.
  • Distributed Scale: Combine tensor, pipeline, and expert parallelism (MoE) across multi-node GPU clusters.
  • Production Benchmarking: Avoid common traps by measuring TTFT, TPOT, and tail latency under realistic workloads.

Stop wasting millions on sub-optimal cloud GPU allocations. Learn how to design, benchmark, and operate multi-tenant, high-throughput, and ultra-low-latency LLM serving architectures today.

Actriz & Políglota
EWA KASP para
Visualizar el vídeo
Ewa Kasp
Libristo tiene la oferta más extensa de literatura en idiomas extranjeros. Por eso compran aquí sus libros.

Sobre el libro

Nombre y apellidos Inference at Full Throttle
Idioma Inglés
Encuadernación Libro - Tapa blanda
Fecha de publicación 2026
Número de páginas 82
EAN 9798192412626
Código Libristo 53520833
Peso 123
Dimensiones 152 x 229 x 4
Regale este libro hoy
Es fácil
1 Añadir al carrito y elegir Entregar como regalo en el checkout 2 Le enviaremos un vale 3 El libro llegará a la dirección del destinatario

Inicio de sesión

Inicie sesión en su cuenta. ¿No tiene una cuenta Libristo? ¡Cree una ahora!

 
obligatorio
obligatorio

¿No tiene cuenta? Descubra las ventajas de tener una cuenta Libristo.

Si tiene una cuenta Libristo, lo tendrá todo bajo control.

Crear una cuenta Libristo
Asesor de libros Libroamiko
Hola, soy Libroamiko, ¿puedo ayudarte?