LIBRISTO
LIBROAMANTO
obligatorio
Entre a formar parte de una comunidad de amantes de los libros del mundo entero y acceda a un sinfín de ventajas. Crear una cuenta gratis
0
Envío gratuito con Zásilkovna para compras superiores a 59.99 €
Mensajería SEUR 4.99 € Mensajería GLS 7.99 € Mensajería Correos 5.49 € Mensajería DHL 5.49 € Punto SEUR 3.99 €

Envío gratis a partir de 69,99 euros.

vLLM Serving

High‑Throughput LLM APIs with PagedAttention and KV Cache Tuning

Idioma InglésInglés
Libro Tapa blanda
Editores NobleTrex Press, junio 2026
"vLLM Serving: High‑Throughput LLM APIs with PagedAttention and KV Cache Tuning"Built for experience... Descripción completa
? points 90 b
36.89 €
Almacenamiento externo Envío en 14-21 días

Hasta 30 días para devoluciones

"vLLM Serving: High‑Throughput LLM APIs with PagedAttention and KV Cache Tuning"

Built for experienced ML systems engineers, platform architects, and performance-minded practitioners, this book is a deep technical guide to serving large language models with vLLM at production scale. Rather than treating inference as a black box, it explains the real control surfaces behind throughput, latency, and memory efficiency. Readers who already know LLM fundamentals but want to reason rigorously about serving behavior will find an internals-first, systems-oriented treatment.

At the core of the book are the mechanisms that make vLLM distinctive: PagedAttention, continuous batching, KV cache design, and scheduler-driven execution. You will learn how request flow, cache allocation, sequence length, prefix reuse, quantized KV storage, and offloading strategies interact to determine concurrency limits and user-visible performance. The book also covers OpenAI-compatible API serving, streaming semantics, realistic benchmarking, and disciplined troubleshooting, so readers can move from conceptual understanding to evidence-based tuning and operational decisions.

The emphasis throughout is on advanced mental models, trade-offs, and production diagnostics rather than introductory walkthroughs. This is a focused guide for readers comfortable with GPU inference, transformer decoding, and performance measurement who want a precise framework for designing, tuning, and operating high-throughput LLM APIs with confidence.

Actriz & Políglota
EWA KASP para
Visualizar el vídeo
Ewa Kasp
Libristo tiene la oferta más extensa de literatura en idiomas extranjeros. Por eso compran aquí sus libros.

Sobre el libro

Nombre y apellidos vLLM Serving
Autor Trex Team
Idioma Inglés
Encuadernación Libro - Tapa blanda
Fecha de publicación 2026
Número de páginas 240
EAN 9798896653417
Código Libristo 54034914
Editores NobleTrex Press
Peso 328
Dimensiones 152 x 229 x 13
Regale este libro hoy
Es fácil
1 Añadir al carrito y elegir Entregar como regalo en el checkout 2 Le enviaremos un vale 3 El libro llegará a la dirección del destinatario

Inicio de sesión

Inicie sesión en su cuenta. ¿No tiene una cuenta Libristo? ¡Cree una ahora!

 
obligatorio
obligatorio

¿No tiene cuenta? Descubra las ventajas de tener una cuenta Libristo.

Si tiene una cuenta Libristo, lo tendrá todo bajo control.

Crear una cuenta Libristo
Asesor de libros Libroamiko
Hola, soy Libroamiko, ¿puedo ayudarte?