Hardware and Chips
About the track
New chips are unlocking new model training and inference capabilities while maneuvering around supply chain constraints. Upgraded TPUs allow 370 tokens per second for Gemini 3.7 Flash, OpenAI is 12x-ing 5.6 Sol speed on Cerebras, and Nvidia is reducing memory on its upcoming Rubin GPU due to rising prices. Hardware is exciting again, at both the startup and enterprise level. This track is for ML engineers looking to go deeper into hardware and kernel knowledge.
Talks cover the new generation of chips, chip architecture, kernels and inference tricks, and recent developments in open source models such as linear attention. Speakers are professionals from chip companies diving into the details of the next generation of hardware. You’ll walk out knowing what the next generation of hardware unlocks in LLMs and other models.
Notes: Confirm the three hardware figures in the second sentence before publishing. Denys wrote the hook with exclamation marks; the words are kept and the punctuation changed to periods to match the QCon register. “manuver” corrected to “maneuvering”.
Track host
Denys Linkov
Head of ML, Wisedocs
See Track Lead Profile
MLOps World | GenAI Summit 2026 is a two days of case studies, workshops, and expo on taking AI/ML and agentic systems into production – at the Etter-Harbin Alumni photo – full-bleed hero or browse files.