Vision Language Models - by Merve Noyan & Andrés Marafioti & Miquel Farré & Orr Zohar (Paperback)

Name: Vision Language Models - by Merve Noyan & Andrés Marafioti & Miquel Farré & Orr Zohar (Paperback)
Brand: O'Reilly Media
SKU: 1008582359
Price: 79.99 USD
Availability: PreOrder

New at

$79.99

Pre-order

Eligible for registries and wish lists

About this item

Highlights

Vision language models (VLMs) combine computer vision and natural language processing to create powerful systems that can interpret, generate, and respond in multimodal contexts.
Author(s): Merve Noyan & Andrés Marafioti & Miquel Farré & Orr Zohar
300 Pages
Computers + Internet, Computer Vision & Pattern Recognition

Description

Book Synopsis

Vision language models (VLMs) combine computer vision and natural language processing to create powerful systems that can interpret, generate, and respond in multimodal contexts. Vision Language Models is a hands-on guide to building real-world VLMs using the most up-to-date stack of machine learning tools from Hugging Face, Meta (PyTorch), NVIDIA (Cuda), OpenAI (CLIP), and others, written by leading researchers and practitioners Merve Noyan, Miquel Farré, Andrés Marafioti, and Orr Zohar. From image captioning and document understanding to advanced zero-shot inference and retrieval-augmented generation (RAG), this book covers the full VLM application and development lifecycle.

Designed for ML engineers, data scientists, and developers, this guide distills cutting-edge VLM research into practical techniques. Readers will learn how to prepare datasets, select the right architectures, fine-tune and deploy models, and apply them to real-world tasks across a range of industries.

Explore core model architectures and alignment techniques
Train and fine-tune VLMs with Hugging Face, PyTorch, and others
Deploy models for applications like image search and captioning
Implement advanced inference strategies, from zero-shot to agentic systems
Build scalable VLM systems ready for production use

Dimensions (Overall): 9.19 Inches (H) x 7.0 Inches (W)

Suggested Age: 22 Years and Up

Number of Pages: 300

Genre: Computers + Internet

Sub-Genre: Computer Vision & Pattern Recognition

Publisher: O'Reilly Media

Format: Paperback

Author: Merve Noyan & Andrés Marafioti & Miquel Farré & Orr Zohar

Language: English

Street Date: September 1, 2026

TCIN: 1008582359

UPC: 9798341624047

Item Number (DPCI): 247-47-5256

Origin: Made in the USA or Imported

If the item details aren’t accurate or complete, we want to know about it.

Shipping details

Estimated ship dimensions: 1 inches length x 7 inches width x 9.19 inches height

Estimated ship weight: 1 pounds

We regret that this item cannot be shipped to PO Boxes.

This item cannot be shipped to the following locations: American Samoa (see also separate entry under AS), Guam (see also separate entry under GU), Northern Mariana Islands, Puerto Rico (see also separate entry under PR), United States Minor Outlying Islands, Virgin Islands, U.S., APO/FPO