All sessions and workshops curated by leading AI/ML practitioners
All Speakers
D. Sculley
D. Sculley
ABOUT THE SPEAKER:
We spend a lot of time thinking about operational issues in AI related to deployment, somewhat less time thinking about adoption, or (dare we say it) acceptance. This talk will touch on some technical pieces in the current AI ops landscape including streaming systems, planning, latency, and on-device models, but most of the time will be spent looking at ways we can move beyond the stale framing of a chatbot, assistant, or customer service agent.
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
We spend a lot of time thinking about operational issues in AI related to deployment, somewhat less time thinking about adoption, or (dare we say it) acceptance. This talk will touch on some technical pieces in the current AI ops landscape including streaming systems, planning, latency, and on-device models, but most of the time will be spent looking at ways we can move beyond the stale framing of a chatbot, assistant, or customer service agent.
Chris Alexiuk
Chris Alexiuk
ABOUT THE SPEAKER:
In this workshop we’ll stand up NemoClaw end to end: install the reference stack, get OpenClaw running inside the OpenShell sandbox, configure inference routing, and lock down a network policy that survives a multi-hour agent session. We’ll walk through the blueprint, the CLI, and the approval flow, then run a real long-lived agent against it and break things on purpose so you know what the layers actually catch.
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
In this workshop we’ll stand up NemoClaw end to end: install the reference stack, get OpenClaw running inside the OpenShell sandbox, configure inference routing, and lock down a network policy that survives a multi-hour agent session. We’ll walk through the blueprint, the CLI, and the approval flow, then run a real long-lived agent against it and break things on purpose so you know what the layers actually catch.
Vashishtha Patil
Vashishtha Patil
ABOUT THE SPEAKER:
Autonomous research agents that propose, implement, and refine ML solutions are priced out by the loop, not the model. Published agents assume a frontier model drives every step, so cost scales with trajectory length on runs that are long by design. Teams either cap the horizon — removing the thing that made it work — or don’t run it at all.
Hypothesis: a small open-weight model runs the loop end to end, with a frontier advisor called in at those points on a metered budget, holding most of the baseline’s result quality at a fraction of its cost. The open question is whether the escalation trigger is reliable enough to justify the calls it buys — our first version often bought advice the loop would have reached on its own.
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
Autonomous research agents that propose, implement, and refine ML solutions are priced out by the loop, not the model. Published agents assume a frontier model drives every step, so cost scales with trajectory length on runs that are long by design. Teams either cap the horizon — removing the thing that made it work — or don’t run it at all.
Hypothesis: a small open-weight model runs the loop end to end, with a frontier advisor called in at those points on a metered budget, holding most of the baseline’s result quality at a fraction of its cost. The open question is whether the escalation trigger is reliable enough to justify the calls it buys — our first version often bought advice the loop would have reached on its own.
Maitrik Patel
Maitrik Patel
ABOUT THE SPEAKER:
Running ML training, AI inference, and agent workflows on three separate orchestration systems is an operational debt that compounds with every new workload type – one system per abstraction means three on-call rotations, three observability stacks, and three onboarding paths. We collapsed this to a single execution layer by designing a type-safe task interface that accommodates non-deterministic agent execution alongside deterministic ML pipelines, backed by container-native isolation and a workload-aware shared scheduler. The unification exposed a hard failure class that static ML infrastructure never encounters: agent tasks are not idempotent, and ML-derived retry semantics caused downstream state corruption in early production until we rebuilt retry logic around explicit checkpointing and action journals. Attendees leave with a four-layer architectural blueprint for unifying these workload types, an honest account of the scheduling contention and retry failures encountered in production, and the specific interface design decisions that made incremental migration rather than forced rewrite the path to adoption.
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
Running ML training, AI inference, and agent workflows on three separate orchestration systems is an operational debt that compounds with every new workload type – one system per abstraction means three on-call rotations, three observability stacks, and three onboarding paths. We collapsed this to a single execution layer by designing a type-safe task interface that accommodates non-deterministic agent execution alongside deterministic ML pipelines, backed by container-native isolation and a workload-aware shared scheduler. The unification exposed a hard failure class that static ML infrastructure never encounters: agent tasks are not idempotent, and ML-derived retry semantics caused downstream state corruption in early production until we rebuilt retry logic around explicit checkpointing and action journals. Attendees leave with a four-layer architectural blueprint for unifying these workload types, an honest account of the scheduling contention and retry failures encountered in production, and the specific interface design decisions that made incremental migration rather than forced rewrite the path to adoption.
Rajiv Shah
Rajiv Shah
ABOUT THE SPEAKER:
You can start a simple agent with a model and a prompt. But you soon realize that improving it requires adjusting what the model can see, what it can do, how it remembers progress, which model handles each step, and what evidence allows the work to stop. Harness engineering is how those pieces become one working system.
This workshop uses coding agents to make that system concrete. Participants will work through six decisions: choosing a harness, designing tools and retrieval, placing context and memory, routing models and reasoning, controlling goals and validation, and deciding when a task benefits from multiple agents. Through hands-on experiments, we will change these settings and inspect what happens in the trace and final result.
Each exercise is paired with an overview of current research on agent harnesses. Together, the research and experiments show how the major components interact and where each approach reaches its limits. Participants will leave with a practical way to reason about the whole agent, not just its model.
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
You can start a simple agent with a model and a prompt. But you soon realize that improving it requires adjusting what the model can see, what it can do, how it remembers progress, which model handles each step, and what evidence allows the work to stop. Harness engineering is how those pieces become one working system.
This workshop uses coding agents to make that system concrete. Participants will work through six decisions: choosing a harness, designing tools and retrieval, placing context and memory, routing models and reasoning, controlling goals and validation, and deciding when a task benefits from multiple agents. Through hands-on experiments, we will change these settings and inspect what happens in the trace and final result.
Each exercise is paired with an overview of current research on agent harnesses. Together, the research and experiments show how the major components interact and where each approach reaches its limits. Participants will leave with a practical way to reason about the whole agent, not just its model.
Kumaran Ponnambalam
Kumaran Ponnambalam
ABOUT THE SPEAKER:
This session examines how a traditionally static enterprise agent workflow was redesigned around AI-driven adaptation. In most enterprise settings, agent behavior is defined through fixed prompts, rules, and workflows that apply broadly across users and tenants. We redesigned that model into an adaptive agent architecture built around explicit policy layers, feedback loops, and governed personalization. Structurally, this changed the system from one-time configuration to continuous learning: policies were separated from model reasoning, policy scope was organized across global, tenant, and user levels, and agent behavior was connected to both explicit and implicit feedback signals that could influence future decisions.
The session will focus on what changed in the system architecture and operating model to support this shift. It will cover how policy-driven personalization was introduced, how feedback became part of the runtime and improvement loop, and how adaptation was managed without losing enterprise control. It will also discuss the practical challenges that emerged when moving from static agents to adaptive ones, including handling conflicting user and tenant needs, keeping personalization aligned with core guardrails, and making behavior changes observable and manageable in production.
Attendees will walk away with a practical understanding of how to redesign enterprise agent workflows around adaptive AI capabilities. They will learn the core concepts behind Personalized Experiential Learning, the structural changes needed to support policy-based adaptation, and the key lessons and best practices for making adaptive agents workable in real enterprise environments. They should leave better equipped to think about how to introduce governed personalization, feedback-driven improvement, and policy evolution into their own AI agent systems.
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
This session examines how a traditionally static enterprise agent workflow was redesigned around AI-driven adaptation. In most enterprise settings, agent behavior is defined through fixed prompts, rules, and workflows that apply broadly across users and tenants. We redesigned that model into an adaptive agent architecture built around explicit policy layers, feedback loops, and governed personalization. Structurally, this changed the system from one-time configuration to continuous learning: policies were separated from model reasoning, policy scope was organized across global, tenant, and user levels, and agent behavior was connected to both explicit and implicit feedback signals that could influence future decisions.
The session will focus on what changed in the system architecture and operating model to support this shift. It will cover how policy-driven personalization was introduced, how feedback became part of the runtime and improvement loop, and how adaptation was managed without losing enterprise control. It will also discuss the practical challenges that emerged when moving from static agents to adaptive ones, including handling conflicting user and tenant needs, keeping personalization aligned with core guardrails, and making behavior changes observable and manageable in production.
Attendees will walk away with a practical understanding of how to redesign enterprise agent workflows around adaptive AI capabilities. They will learn the core concepts behind Personalized Experiential Learning, the structural changes needed to support policy-based adaptation, and the key lessons and best practices for making adaptive agents workable in real enterprise environments. They should leave better equipped to think about how to introduce governed personalization, feedback-driven improvement, and policy evolution into their own AI agent systems.
Balaji Varadarajan
Balaji Varadarajan
ABOUT THE SPEAKER:
We will start with two diagnostic tools: Little’s Law which translates your QPS and target latency into the concurrency and batch size your system actually needs and the roofline model which tells you whether prefill (compute-bound) or decode (memory-bound) is your real constraint before you reach for a fix..
Will walk through the levers that move the needle most – MoE vs. dense architecture and what that does to your communication pattern FP8/INT4 quantization and where it costs you accuracy, TP/EP/DP parallelism and when each is the right tool and batching to the B_sat knee before compute-bound latency hockey-sticks..
Will cover concrete tuning playbooks for two workload types:
- latency-sensitive traffic like chat and agents where you’re optimizing TTFT and P99 ITL with chunked prefill and modest batch sizes
- Throughput-sensitive traffic like batch summarization and evals where you push past B_sat and let queuing work in your favor.
Most production systems live in both worlds at once, so we’ll need a balanced configuration to tackle both.
You will leave with a 7-step decision framework – characterize your workload, pick your model and quantization, benchmark on candidate hardware, find the knee, size your deployment with Little’s Law, calculate total cost of ownership, and plan for autoscaling that you can apply to your own inference stack..
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
We will start with two diagnostic tools: Little’s Law which translates your QPS and target latency into the concurrency and batch size your system actually needs and the roofline model which tells you whether prefill (compute-bound) or decode (memory-bound) is your real constraint before you reach for a fix..
Will walk through the levers that move the needle most – MoE vs. dense architecture and what that does to your communication pattern FP8/INT4 quantization and where it costs you accuracy, TP/EP/DP parallelism and when each is the right tool and batching to the B_sat knee before compute-bound latency hockey-sticks..
Will cover concrete tuning playbooks for two workload types:
- latency-sensitive traffic like chat and agents where you’re optimizing TTFT and P99 ITL with chunked prefill and modest batch sizes
- Throughput-sensitive traffic like batch summarization and evals where you push past B_sat and let queuing work in your favor.
Most production systems live in both worlds at once, so we’ll need a balanced configuration to tackle both.
You will leave with a 7-step decision framework – characterize your workload, pick your model and quantization, benchmark on candidate hardware, find the knee, size your deployment with Little’s Law, calculate total cost of ownership, and plan for autoscaling that you can apply to your own inference stack..
Anne Griffin
Anne Griffin
ABOUT THE SPEAKER:
When Hugging Face went to investigate this summer’s OpenAI containment breach, the commercial models they reached for refused to help, because their safety guardrails couldn’t tell defending from attacking apart. Instead, Hugging Face used an open weight model on their own infrastructure and finished the forensics in hours with none of the attacker data leaving their environment. It’s the perfect example of how open weight models let you control your guardrails, privacy, governance, latency, and reliability.
More companies are asking if they should use and self host open weight models, and this talk will walk through a framework to determine when it does and doesn’t make sense. By the end of this talk, attendees will understand what the benefits and drawbacks are of open weight models, and how they can bring both technical and business value.
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
When Hugging Face went to investigate this summer’s OpenAI containment breach, the commercial models they reached for refused to help, because their safety guardrails couldn’t tell defending from attacking apart. Instead, Hugging Face used an open weight model on their own infrastructure and finished the forensics in hours with none of the attacker data leaving their environment. It’s the perfect example of how open weight models let you control your guardrails, privacy, governance, latency, and reliability.
More companies are asking if they should use and self host open weight models, and this talk will walk through a framework to determine when it does and doesn’t make sense. By the end of this talk, attendees will understand what the benefits and drawbacks are of open weight models, and how they can bring both technical and business value.
Fuzail Khan
Fuzail Khan
ABOUT THE SPEAKER:
The ranking and recommendation systems landscape is being transformed in the generative era. This talk reports on the experience of building and shipping the training infrastructure behind the first generative recommender in production at Meta, covering both stages of the recipe: pre-training to acquire the generative capability and post-training RL to align generation with the ranking objective.
Generative recommenders bring distinct challenges to end-to-end performance and scalability. This primarily arises from a mixed architecture that consists of both recommender-native large sparse embedding tables and LLM-based decoders. This then generates semantic IDs translating to real-world use cases in production such as finding the right advertisement for a given user. We inherit the communication profile of a sparse recommender as well as the autoregressive nature of a large language model for which an end-to-end systems blueprint simply does not exist.
For pre-training, we make the significant change from discriminative to generative recommendation. We talk about feature processing, data loading and the user-modeling path while building the LLM decoder and semantic ID tokenization on top while addressing real-world productionization challenges and end-to-end LLM performance analysis.
Reinforcement learning was implemented as an extensible framework where rewards, losses, RL algorithms, reference models and generation strategies are pluggable. We’ll walk through the reward design and the end-to-end post-training pipeline that prioritizes scalability at the production scale at Meta.
We’ll close with end-to-end optimization work – redundant computation elimination, specialized kernels for very short sequences, dense-sparse pipelining and hybrid embedding placement that satisfies large-scale performance requirements in compute and throughput.
We believe the key output from this session is leaving attendees with a crisp understanding of generative recommendation systems in large-scale production. We want to ensure we stay practical and address real-world production challenges that are relevant to developers and builders in AI infrastructure today. We do this by diving into the technicalities of enabling a large-scale generative retrieval system as well as concrete performance optimizations that enable the shift from a research prototype to a highly-optimized production system critical to revenue. We also feel the training systems and infrastructure-side of LLMs in production is rarely addressed at least relative to model architecture and quality, and we hope to fill that gap with this session.
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
The ranking and recommendation systems landscape is being transformed in the generative era. This talk reports on the experience of building and shipping the training infrastructure behind the first generative recommender in production at Meta, covering both stages of the recipe: pre-training to acquire the generative capability and post-training RL to align generation with the ranking objective.
Generative recommenders bring distinct challenges to end-to-end performance and scalability. This primarily arises from a mixed architecture that consists of both recommender-native large sparse embedding tables and LLM-based decoders. This then generates semantic IDs translating to real-world use cases in production such as finding the right advertisement for a given user. We inherit the communication profile of a sparse recommender as well as the autoregressive nature of a large language model for which an end-to-end systems blueprint simply does not exist.
For pre-training, we make the significant change from discriminative to generative recommendation. We talk about feature processing, data loading and the user-modeling path while building the LLM decoder and semantic ID tokenization on top while addressing real-world productionization challenges and end-to-end LLM performance analysis.
Reinforcement learning was implemented as an extensible framework where rewards, losses, RL algorithms, reference models and generation strategies are pluggable. We’ll walk through the reward design and the end-to-end post-training pipeline that prioritizes scalability at the production scale at Meta.
We’ll close with end-to-end optimization work – redundant computation elimination, specialized kernels for very short sequences, dense-sparse pipelining and hybrid embedding placement that satisfies large-scale performance requirements in compute and throughput.
We believe the key output from this session is leaving attendees with a crisp understanding of generative recommendation systems in large-scale production. We want to ensure we stay practical and address real-world production challenges that are relevant to developers and builders in AI infrastructure today. We do this by diving into the technicalities of enabling a large-scale generative retrieval system as well as concrete performance optimizations that enable the shift from a research prototype to a highly-optimized production system critical to revenue. We also feel the training systems and infrastructure-side of LLMs in production is rarely addressed at least relative to model architecture and quality, and we hope to fill that gap with this session.
Sandeep Bharadwaj Mannapur
Sandeep Bharadwaj Mannapur
ABOUT THE SPEAKER:
We caught our LLM quality problem the wrong way: through customer complaints. Three weeks of degraded responses had already shipped. Every monitoring signal we had was green the entire time because we were measuring the wrong things. Latency, error rate, embedding similarity: none of them captured what was actually happening, which was that response quality had quietly gotten worse in ways users noticed but our dashboards could not.
After that incident we rebuilt how we think about LLM observability. The core insight was that output quality is not the same as system health, and you cannot infer one from the other. We needed a separate signal layer built from production behavior: how users responded to answers, where downstream tasks broke down, where retrieval and response stopped agreeing with each other.
Getting this right took a few tries. A naive implementation fires constantly on normal LLM output variance. The real work was designing alerts that distinguish actual degradation from noise, and calibrating thresholds against real incident history rather than theoretical bounds.
You will leave with a taxonomy of LLM drift types and how each one shows up differently in production, the behavioral signals that actually correlate with quality degradation, and an alert pattern that catches real drift early without burying your team in false positives.
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
We caught our LLM quality problem the wrong way: through customer complaints. Three weeks of degraded responses had already shipped. Every monitoring signal we had was green the entire time because we were measuring the wrong things. Latency, error rate, embedding similarity: none of them captured what was actually happening, which was that response quality had quietly gotten worse in ways users noticed but our dashboards could not.
After that incident we rebuilt how we think about LLM observability. The core insight was that output quality is not the same as system health, and you cannot infer one from the other. We needed a separate signal layer built from production behavior: how users responded to answers, where downstream tasks broke down, where retrieval and response stopped agreeing with each other.
Getting this right took a few tries. A naive implementation fires constantly on normal LLM output variance. The real work was designing alerts that distinguish actual degradation from noise, and calibrating thresholds against real incident history rather than theoretical bounds.
You will leave with a taxonomy of LLM drift types and how each one shows up differently in production, the behavioral signals that actually correlate with quality degradation, and an alert pattern that catches real drift early without burying your team in false positives.
Drew Crawford
Drew Crawford
ABOUT THE SPEAKER:
Somewhere in a datacenter right now, one frontier model is telling another that it’s the town doctor. It is not the town doctor. It killed someone four turns ago, and it’s about to get a third model lynched for it.
Most LLM benchmarks are exams: a model alone in a room with a test paper. Mafia is a room where some of the agents are lying, everyone knows some of them are lying, and the game is figuring out which. It demands recursive theory of mind, deception that stays consistent under adversarial re-reading, and long-context discipline — and it punishes output indiscipline like no static eval can: a model that rambles or breaks format doesn’t get a bad grade, it gets voted out. The other players are the evaluation harness.
This talk covers what it took to run AI Bot Mafia in production: 40,000 lines of Rust; the streaming pathologies of 104 model configurations (chunks that split UTF-8 characters mid-byte, keepalive-only hangs no single timeout catches, malformed tool calls that poison history); and cost governance where prompt-cache economics made cache locality a game-design constraint, enforced by spend fuses and a concurrency queue.
Then the leaderboard: why win rates and Elo are statistically indefensible here, and what works instead — ridge logistic regression over per-seat signed features, bootstrapped 200×, ranked by the p10 lower bound, with a machine-readable reliability state that says “insufficient data” out loud.
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
Somewhere in a datacenter right now, one frontier model is telling another that it’s the town doctor. It is not the town doctor. It killed someone four turns ago, and it’s about to get a third model lynched for it.
Most LLM benchmarks are exams: a model alone in a room with a test paper. Mafia is a room where some of the agents are lying, everyone knows some of them are lying, and the game is figuring out which. It demands recursive theory of mind, deception that stays consistent under adversarial re-reading, and long-context discipline — and it punishes output indiscipline like no static eval can: a model that rambles or breaks format doesn’t get a bad grade, it gets voted out. The other players are the evaluation harness.
This talk covers what it took to run AI Bot Mafia in production: 40,000 lines of Rust; the streaming pathologies of 104 model configurations (chunks that split UTF-8 characters mid-byte, keepalive-only hangs no single timeout catches, malformed tool calls that poison history); and cost governance where prompt-cache economics made cache locality a game-design constraint, enforced by spend fuses and a concurrency queue.
Then the leaderboard: why win rates and Elo are statistically indefensible here, and what works instead — ridge logistic regression over per-seat signed features, bootstrapped 200×, ranked by the p10 lower bound, with a machine-readable reliability state that says “insufficient data” out loud.
Ishaan Sehgal
Ishaan Sehgal
ABOUT THE SPEAKER:
At Omnara, long-running agent sessions exposed a failure mode hidden by short-lived agents: the agent’s state was coupled to the process, model provider, or runtime executing it. When a worker crashed, a client disconnected, or execution moved to another machine, recovering the agent’s work became brittle.
We redesigned agent execution around an append-only session log as the source of truth. Model outputs, tool calls and results, user interventions, and other state transitions are persisted as ordered events; the live agent state becomes a replayable projection of that history.
This structural change decouples agent identity from any one model, worker, or harness. It enables crash recovery, resumability, branching, real-time observability, and migration across models and machines, but introduces difficult systems questions around ordering, idempotency, partial tool execution, checkpointing, compaction, concurrent writers, and log ownership.
This session walks through the production failures that motivated the redesign, the resulting event model and recovery path, and the tradeoffs we encountered. Using production traces and forced-failure examples, we will show what happens when a worker dies mid-tool call, how a session resumes under a different runtime, and how replay costs change as histories grow.
Attendees will leave with a practical blueprint for deciding what belongs in the durable log, what can be recomputed, and how to build agent sessions that survive process, machine, and model failure without locking continuity to a single provider.
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
At Omnara, long-running agent sessions exposed a failure mode hidden by short-lived agents: the agent’s state was coupled to the process, model provider, or runtime executing it. When a worker crashed, a client disconnected, or execution moved to another machine, recovering the agent’s work became brittle.
We redesigned agent execution around an append-only session log as the source of truth. Model outputs, tool calls and results, user interventions, and other state transitions are persisted as ordered events; the live agent state becomes a replayable projection of that history.
This structural change decouples agent identity from any one model, worker, or harness. It enables crash recovery, resumability, branching, real-time observability, and migration across models and machines, but introduces difficult systems questions around ordering, idempotency, partial tool execution, checkpointing, compaction, concurrent writers, and log ownership.
This session walks through the production failures that motivated the redesign, the resulting event model and recovery path, and the tradeoffs we encountered. Using production traces and forced-failure examples, we will show what happens when a worker dies mid-tool call, how a session resumes under a different runtime, and how replay costs change as histories grow.
Attendees will leave with a practical blueprint for deciding what belongs in the durable log, what can be recomputed, and how to build agent sessions that survive process, machine, and model failure without locking continuity to a single provider.
Amit Kumar Padhy
Amit Kumar Padhy
ABOUT THE SPEAKER:
Modern commerce platforms don’t fail because of missing features, they fail at the seams.
A product is created in Catalog, but pricing is incomplete. Promotions don’t qualify. Tax blocks specific regions. Localization lags. The system says “”launched,” but the business knows it isn’t. These aren’t edge cases, they are the steady state of distributed commerce.
This session replaces traditional workflow orchestration with a multi-agent, swarm-based execution model powered by LLMs and agentic AI, coordinating Pricing, Catalog, Promotions, Tax, and Compliance in real time.
Agent Roles. Planner Agents decompose onboarding goals into executable plans using ReAct-style tool-aware reasoning and function calling. Domain Agents for Pricing, Catalog, and Compliance execute directly against APIs, Pricing Runtime, Offer Systems, Billing Preview, Tax engines. Validator Agents enforce policy rules, regional compliance, and pricing integrity at every step. A Coordinator Agent maintains shared state and resolves cross-domain conflicts through blackboard-pattern memory.
Execution Model. Agents subscribe to domain events, SKU created, price missing, policy violation, via streaming infrastructure. LLM-backed planners generate dynamic execution graphs, not fixed DAGs. Each agent invokes APIs as tools, binding service calls to reasoning steps. Coordination flows through a shared state store and event-driven updates, enabling adaptive planning under partial data, cross-service dependency resolution, and parallel execution with conflict detection.
Failure Handling. No blind retries. LLM-guided compensations roll back partial pricing, recompute discounts, and re-trigger downstream syncs based on causal analysis. No opaque errors. Every decision produces a chain-of-thought trace with tool invocation logs and reasoning outputs. No brittle state. Recovery is idempotent and event-sourced.
Onboarding Flow. Catalog Agents handle onboarding for products by ensuring the product definition is created, and the Pricing agent ensures SKUs are priced. Promotion Agents conduct market checks, provide qualified promotions and their onboarding, and resolve discounts using a rule-based and LLM-hybrid evaluation.
Compliance Agents validate regulatory constraints per geography. Localization Agents assess market-readiness signals, and product merchandising is localized accordingly. The Coordinator drives convergence toward “sellable state”.
Engineering Stack. Kafka-backed event contracts for agent triggers. Function-calling LLM agents with API abstraction layers. Vector DB plus state store for context propagation. Human-in-the-loop decision gates only at high-risk boundaries. Immutable execution logs tracing every action and compensation. Observability is built around decision reasoning, not just service health.
Takeaway. A working blueprint for LLM-powered agent swarms that handle uncertainty across distributed commerce, recover through intelligent compensation, and ensure products ship globally, not just deploy.
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
Modern commerce platforms don’t fail because of missing features, they fail at the seams.
A product is created in Catalog, but pricing is incomplete. Promotions don’t qualify. Tax blocks specific regions. Localization lags. The system says “”launched,” but the business knows it isn’t. These aren’t edge cases, they are the steady state of distributed commerce.
This session replaces traditional workflow orchestration with a multi-agent, swarm-based execution model powered by LLMs and agentic AI, coordinating Pricing, Catalog, Promotions, Tax, and Compliance in real time.
Agent Roles. Planner Agents decompose onboarding goals into executable plans using ReAct-style tool-aware reasoning and function calling. Domain Agents for Pricing, Catalog, and Compliance execute directly against APIs, Pricing Runtime, Offer Systems, Billing Preview, Tax engines. Validator Agents enforce policy rules, regional compliance, and pricing integrity at every step. A Coordinator Agent maintains shared state and resolves cross-domain conflicts through blackboard-pattern memory.
Execution Model. Agents subscribe to domain events, SKU created, price missing, policy violation, via streaming infrastructure. LLM-backed planners generate dynamic execution graphs, not fixed DAGs. Each agent invokes APIs as tools, binding service calls to reasoning steps. Coordination flows through a shared state store and event-driven updates, enabling adaptive planning under partial data, cross-service dependency resolution, and parallel execution with conflict detection.
Failure Handling. No blind retries. LLM-guided compensations roll back partial pricing, recompute discounts, and re-trigger downstream syncs based on causal analysis. No opaque errors. Every decision produces a chain-of-thought trace with tool invocation logs and reasoning outputs. No brittle state. Recovery is idempotent and event-sourced.
Onboarding Flow. Catalog Agents handle onboarding for products by ensuring the product definition is created, and the Pricing agent ensures SKUs are priced. Promotion Agents conduct market checks, provide qualified promotions and their onboarding, and resolve discounts using a rule-based and LLM-hybrid evaluation.
Compliance Agents validate regulatory constraints per geography. Localization Agents assess market-readiness signals, and product merchandising is localized accordingly. The Coordinator drives convergence toward “sellable state”.
Engineering Stack. Kafka-backed event contracts for agent triggers. Function-calling LLM agents with API abstraction layers. Vector DB plus state store for context propagation. Human-in-the-loop decision gates only at high-risk boundaries. Immutable execution logs tracing every action and compensation. Observability is built around decision reasoning, not just service health.
Takeaway. A working blueprint for LLM-powered agent swarms that handle uncertainty across distributed commerce, recover through intelligent compensation, and ensure products ship globally, not just deploy.
Vishvesh Pandey
Vishvesh Pandey
ABOUT THE SPEAKER:
Most practitioners instrumenting AI agents in production analytics workflows start by tracking token consumption — and discover too late that token spend tells you what you spent, not whether it was worth it. This talk walks through the operational reframe underway in financial services and adjacent regulated industries: from cost-per-token to cost-per-resolved-task as the primary economic unit, with call-level attribution as the foundation that makes everything else possible.
I’ll cover four cost engineering patterns that hold up in production analytics environments: (1) instrumentation at the call level — what to log, where it lives, who owns it; (2) model tiering and routing — when a frontier model earns its premium and when a smaller fine-tuned model wins; (3) the output-token asymmetry that makes structured outputs and chain-of-thought hygiene more economically consequential than they appear; and (4) cost containment as an engineering practice with runtime guardrails, not a finance review after the fact.
I’ll also be honest about what doesn’t generalize: caching strategies that worked for one workload and broke another, routing logic that added more latency than it saved, and the political dimension of asking analysts to justify their agent spend. Attendees will leave with a maturity model for agent cost engineering, a concrete instrumentation checklist they can apply next week, and a vocabulary for the cost conversation that bridges engineering and finance — because in regulated industries, that’s the conversation that decides whether agents scale or stall.
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
Most practitioners instrumenting AI agents in production analytics workflows start by tracking token consumption — and discover too late that token spend tells you what you spent, not whether it was worth it. This talk walks through the operational reframe underway in financial services and adjacent regulated industries: from cost-per-token to cost-per-resolved-task as the primary economic unit, with call-level attribution as the foundation that makes everything else possible.
I’ll cover four cost engineering patterns that hold up in production analytics environments: (1) instrumentation at the call level — what to log, where it lives, who owns it; (2) model tiering and routing — when a frontier model earns its premium and when a smaller fine-tuned model wins; (3) the output-token asymmetry that makes structured outputs and chain-of-thought hygiene more economically consequential than they appear; and (4) cost containment as an engineering practice with runtime guardrails, not a finance review after the fact.
I’ll also be honest about what doesn’t generalize: caching strategies that worked for one workload and broke another, routing logic that added more latency than it saved, and the political dimension of asking analysts to justify their agent spend. Attendees will leave with a maturity model for agent cost engineering, a concrete instrumentation checklist they can apply next week, and a vocabulary for the cost conversation that bridges engineering and finance — because in regulated industries, that’s the conversation that decides whether agents scale or stall.
Tony Blank
Tony Blank
ABOUT THE SPEAKER:
We stopped equipping humans and started automating their worst hour. Instead of teaching people to fish, we built the pond: precise background agents pointed at repeated, skill-independent time-sinks where ROI was measurable on day one. The first build was a bug-triage agent wired to GitHub, Jira, and Notion; every agent that followed was measured the same way — Time Saved Per Task (TSPT) per week — against an amortized cost model (build + inference + maintenance vs. hours saved × loaded hourly cost).
Structurally, the change was moving from a portfolio of per-person tool subscriptions to a portfolio of scoped, measurable background agents with a pre-build attribution conversation baked in. The decision rubric — Sharpness × Cadence × Skill Tax → Pond / Boat / Net / Wade — tells you when to automate fully, when to augment, when to keep a human in the loop, and when the math doesn’t work yet.
Attendees will leave able to: score their own candidate workflows on that rubric, instrument TSPT, build a defensible amortization story, and avoid the most expensive failure mode in AI enablement — “give everyone the tools” — by building the pond first.
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
We stopped equipping humans and started automating their worst hour. Instead of teaching people to fish, we built the pond: precise background agents pointed at repeated, skill-independent time-sinks where ROI was measurable on day one. The first build was a bug-triage agent wired to GitHub, Jira, and Notion; every agent that followed was measured the same way — Time Saved Per Task (TSPT) per week — against an amortized cost model (build + inference + maintenance vs. hours saved × loaded hourly cost).
Structurally, the change was moving from a portfolio of per-person tool subscriptions to a portfolio of scoped, measurable background agents with a pre-build attribution conversation baked in. The decision rubric — Sharpness × Cadence × Skill Tax → Pond / Boat / Net / Wade — tells you when to automate fully, when to augment, when to keep a human in the loop, and when the math doesn’t work yet.
Attendees will leave able to: score their own candidate workflows on that rubric, instrument TSPT, build a defensible amortization story, and avoid the most expensive failure mode in AI enablement — “give everyone the tools” — by building the pond first.
Robert Lewis
Robert Lewis
ABOUT THE SPEAKER:
Large language model (LLM) agents can now draft executable business logic — compliance rules, audit checks, detection heuristics — orders of magnitude faster than human engineers. In regulated domains, however, no such logic ships without a human domain expert’s sign-off. We identify a subtle and, in our experience, pervasive failure mode in this human-in-the-loop pattern: review/serving skew, in which the evidence a domain expert reviews is generated by a re-implementation of the candidate logic (a survey script, a notebook, a prompt-level approximation) rather than by the production engine that will ultimately execute it. When the two implementations diverge — and they reliably do — expert approval validates the wrong behavior. We present a discipline that eliminates this failure mode by construction: candidate rules are drafted directly into the production repository in an inactive-but-runnable state, all review evidence is generated by the production execution path, refinement edits the production artifact itself, and activation is reduced to a one-line data change. We describe the discipline in the context of a deployed multi-agent system that authors, validates, and ships declarative compliance rules for automated audit of collision-repair insurance estimates, evaluated against a corpus of approximately 497,000 production documents. The system couples eight specialized LLM agents with deterministic safety rails — golden-fixture regression, an activation allowlist, single-rule change isolation, and corpus-governance constraints — and drives refinement through a bounded, measurable convergence loop against expert false-positive verdicts. We formalize the discipline as five principles, catalog the anti-patterns it forbids, report operational experience from sixteen production rules, and argue that the pattern transfers directly to clinical decision support, financial transaction monitoring, legal document review, and content moderation.
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
Large language model (LLM) agents can now draft executable business logic — compliance rules, audit checks, detection heuristics — orders of magnitude faster than human engineers. In regulated domains, however, no such logic ships without a human domain expert’s sign-off. We identify a subtle and, in our experience, pervasive failure mode in this human-in-the-loop pattern: review/serving skew, in which the evidence a domain expert reviews is generated by a re-implementation of the candidate logic (a survey script, a notebook, a prompt-level approximation) rather than by the production engine that will ultimately execute it. When the two implementations diverge — and they reliably do — expert approval validates the wrong behavior. We present a discipline that eliminates this failure mode by construction: candidate rules are drafted directly into the production repository in an inactive-but-runnable state, all review evidence is generated by the production execution path, refinement edits the production artifact itself, and activation is reduced to a one-line data change. We describe the discipline in the context of a deployed multi-agent system that authors, validates, and ships declarative compliance rules for automated audit of collision-repair insurance estimates, evaluated against a corpus of approximately 497,000 production documents. The system couples eight specialized LLM agents with deterministic safety rails — golden-fixture regression, an activation allowlist, single-rule change isolation, and corpus-governance constraints — and drives refinement through a bounded, measurable convergence loop against expert false-positive verdicts. We formalize the discipline as five principles, catalog the anti-patterns it forbids, report operational experience from sixteen production rules, and argue that the pattern transfers directly to clinical decision support, financial transaction monitoring, legal document review, and content moderation.
WHAT YOU’LL LEARN:
Robert Joel Lewis is a computational social scientist and senior AI engineer whose work bridges artificial intelligence, statistics, software development, and human behavior. Across a career spanning academia and industry, he has published more than 30 scholarly works and built AI-driven products for ecommerce, entertainment, mobile gaming, and collision repair. His research examines entertainment as an attention-capturing technology, including a forthcoming article in the Journal of Media Psychology. Based in Austin, Texas, he specializes in turning complex, messy problems into rigorous systems that people can actually use.
PREREQUISITE KNOWLEDGE
Poonam Lamba
Poonam Lamba
ABOUT THE SPEAKER:
We redesigned distributed GPU orchestration for RL post-training and batch inference in the open-source llm-d platform. Structurally, we replaced static GPU/TPU locking with a three-tier co-operative time-slicing system:
- Application Layer: Workloads signal phase boundaries (rollouts, training, batch inference) via explicit acquire() and yield() APIs.
- Cluster Orchestrator: Manages lock queues to dynamically interleave complementary jobs onto shared hardware during idle phases.
- Node Snapshot Agent: Executes fast sub-second state swaps between GPU/TPU VRAM and host DRAM, enabling instant context switching without container restarts.
Attendees will walk away with:
- Drive 70%+ GPU/TPU Utilization: Understand how time-slicing reclaims idle hardware during RL loops and batch inference without impacting convergence.
- Architect Rapid Memory Swapping: Apply VRAM-to-DRAM snapshotting strategies for ultra-fast GPU/TPU context switching.
- Deploy on Kubernetes: Configure llm-d and K8s orchestrators to interleave RL and batch inference on shared clusters.
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
We redesigned distributed GPU orchestration for RL post-training and batch inference in the open-source llm-d platform. Structurally, we replaced static GPU/TPU locking with a three-tier co-operative time-slicing system:
- Application Layer: Workloads signal phase boundaries (rollouts, training, batch inference) via explicit acquire() and yield() APIs.
- Cluster Orchestrator: Manages lock queues to dynamically interleave complementary jobs onto shared hardware during idle phases.
- Node Snapshot Agent: Executes fast sub-second state swaps between GPU/TPU VRAM and host DRAM, enabling instant context switching without container restarts.
Attendees will walk away with:
- Drive 70%+ GPU/TPU Utilization: Understand how time-slicing reclaims idle hardware during RL loops and batch inference without impacting convergence.
- Architect Rapid Memory Swapping: Apply VRAM-to-DRAM snapshotting strategies for ultra-fast GPU/TPU context switching.
- Deploy on Kubernetes: Configure llm-d and K8s orchestrators to interleave RL and batch inference on shared clusters.
Kalpesh Sutaria
Kalpesh Sutaria
ABOUT THE SPEAKER:
A leaderboard win is not a product. Nemotron Retriever’s models rank #1 on RTEB — but topping a benchmark and running efficiently inside a customer’s production are two very different problems. This talk is the engineering story of closing that gap.
I’ll walk through our journey rebuilding the Retriever inference stack in Rust with a single obsession: treating state-of-the-art performance as a first-class product requirement, not a post-hoc optimization. We’ll get concrete about the decisions that mattered — what we measured, where we spent effort, the throughput, memory, and footprint wins we chased, and how dramatic efficiency gains unlocked deployment scenarios (including edge and on-device) that simply weren’t possible before. I’ll also share what didn’t work, the tension between shipping fast and building durable, and how we leaned on AI-assisted development to move faster than the roadmap assumed.
You’ll leave with a mental model for treating inference performance as a product and go-to-market lever — and a practical playbook for taking research-grade models to production without leaving speed, cost, or reach on the table.
Key takeaways:
- Why “performance is the product”: efficiency is an adoption and go-to-market lever, not a cost center.
- Concrete decisions from a Rust-based inference rebuild — what to measure, where to optimize, and how to know it worked.
- How latency and footprint reductions open entirely new deployment surfaces (edge, on-device) for the same models.
- Navigating the research-to-production seam: turning SoTA benchmark results into reliable, cheap, fast serving customers can bet on.
Who should attend / level? AI engineers, ML platform/infrastructure engineers, and engineering leaders who build, deploy, or operate inference and serving systems. Intermediate.
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
A leaderboard win is not a product. Nemotron Retriever’s models rank #1 on RTEB — but topping a benchmark and running efficiently inside a customer’s production are two very different problems. This talk is the engineering story of closing that gap.
I’ll walk through our journey rebuilding the Retriever inference stack in Rust with a single obsession: treating state-of-the-art performance as a first-class product requirement, not a post-hoc optimization. We’ll get concrete about the decisions that mattered — what we measured, where we spent effort, the throughput, memory, and footprint wins we chased, and how dramatic efficiency gains unlocked deployment scenarios (including edge and on-device) that simply weren’t possible before. I’ll also share what didn’t work, the tension between shipping fast and building durable, and how we leaned on AI-assisted development to move faster than the roadmap assumed.
You’ll leave with a mental model for treating inference performance as a product and go-to-market lever — and a practical playbook for taking research-grade models to production without leaving speed, cost, or reach on the table.
Key takeaways:
- Why “performance is the product”: efficiency is an adoption and go-to-market lever, not a cost center.
- Concrete decisions from a Rust-based inference rebuild — what to measure, where to optimize, and how to know it worked.
- How latency and footprint reductions open entirely new deployment surfaces (edge, on-device) for the same models.
- Navigating the research-to-production seam: turning SoTA benchmark results into reliable, cheap, fast serving customers can bet on.
Who should attend / level? AI engineers, ML platform/infrastructure engineers, and engineering leaders who build, deploy, or operate inference and serving systems. Intermediate.
Nadia Rauch
Nadia Rauch
ABOUT THE SPEAKER:
We built a multi-agent pipeline to automate a complex, document-intensive enterprise workflow. The initial system ran sequentially through a chain of specialized agent roles, each handling a distinct task in the process, using a single frontier model throughout. It worked. It was also slow and expensive, and we didn’t know why.
Rather than optimize blindly, we profiled first. The results were not where we expected: 67% of total latency came from a small minority of the agent roles, and the primary bottleneck was not the model — it was sequential chaining: independent work was being processed one step at a time instead of concurrently. Identifying and parallelizing the roles with no inter-dependency reduced end-to-end runtime by roughly 40%.*
The second experiment compared model tiers (frontier, mid-tier, lightweight) across each agent role independently, measuring accuracy, latency, and cost per role. The finding cuts against the default assumption: the roles that appeared most cognitively demanding required frontier models, but the accuracy gap was smaller than expected. The roles where model downgrade failed were the ones responsible for precise structured extraction — tasks where errors propagate silently downstream and surface only at the output. Swapping to mid-tier models on roles that tolerated it reduced per-run cost by ~40%, with an estimated <2% accuracy impact at the role level and <1% at the pipeline output level.*
The third experiment compared orchestration harnesses — evaluating how the choice of agentic framework affects runtime overhead, observability, and the ease of implementing the parallelization and model-swap changes described above. Framework choice turned out to matter more than expected for operational concerns: debugging multi-agent failures, tracing costs per agent, and modifying execution flow without rewriting pipeline logic.
Operating in a regulated enterprise environment added one constraint worth naming: every agent decision needs to be auditable. This ruled out certain optimization shortcuts that would have been acceptable in other contexts and shaped how we defined “accurate enough” per role.
This talk covers the profiling methodology, the per-role model comparison framework, the parallelization decisions, the harness comparison, and the operational lessons — including what we’d instrument from day one if we rebuilt the system today.
*Cost and accuracy figures for the parallelization and model-tier experiments are preliminary; the formal evaluation is in progress and will be updated with measured results before the talk.
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
We built a multi-agent pipeline to automate a complex, document-intensive enterprise workflow. The initial system ran sequentially through a chain of specialized agent roles, each handling a distinct task in the process, using a single frontier model throughout. It worked. It was also slow and expensive, and we didn’t know why.
Rather than optimize blindly, we profiled first. The results were not where we expected: 67% of total latency came from a small minority of the agent roles, and the primary bottleneck was not the model — it was sequential chaining: independent work was being processed one step at a time instead of concurrently. Identifying and parallelizing the roles with no inter-dependency reduced end-to-end runtime by roughly 40%.*
The second experiment compared model tiers (frontier, mid-tier, lightweight) across each agent role independently, measuring accuracy, latency, and cost per role. The finding cuts against the default assumption: the roles that appeared most cognitively demanding required frontier models, but the accuracy gap was smaller than expected. The roles where model downgrade failed were the ones responsible for precise structured extraction — tasks where errors propagate silently downstream and surface only at the output. Swapping to mid-tier models on roles that tolerated it reduced per-run cost by ~40%, with an estimated <2% accuracy impact at the role level and <1% at the pipeline output level.*
The third experiment compared orchestration harnesses — evaluating how the choice of agentic framework affects runtime overhead, observability, and the ease of implementing the parallelization and model-swap changes described above. Framework choice turned out to matter more than expected for operational concerns: debugging multi-agent failures, tracing costs per agent, and modifying execution flow without rewriting pipeline logic.
Operating in a regulated enterprise environment added one constraint worth naming: every agent decision needs to be auditable. This ruled out certain optimization shortcuts that would have been acceptable in other contexts and shaped how we defined “accurate enough” per role.
This talk covers the profiling methodology, the per-role model comparison framework, the parallelization decisions, the harness comparison, and the operational lessons — including what we’d instrument from day one if we rebuilt the system today.
*Cost and accuracy figures for the parallelization and model-tier experiments are preliminary; the formal evaluation is in progress and will be updated with measured results before the talk.
Aryan Dhar
Aryan Dhar
ABOUT THE SPEAKER:
This talk presents a practical framework for evaluating increasingly complex, non-deterministic AI systems, moving from individual models to complete production pipelines. I begin with a page-stream segmentation model, where standard metrics concealed important failure modes at Wisedocs. For example, we discovered that the accuracy of identifying a document boundary is not the same as the ability to reconstruct the complete document span. Furthemore, errors on some document types carried greater downstream risk than others, and small aggregate improvements did not necessarily translate into meaningful business outcomes. This case study shows how evaluation for us evolved towards task-specific metrics, sensitive failure slices, strict regression testing, and more business-facing metrics of reliability.
I then move to LLM-based systems using Medical Long-Context Reasoning, or MLCR, as a case study. This was a benchmark we developed at Wisedocs earlier this summer. I will examine how to evaluate systems such as agentic chatbots and AI insight-generation services when outputs are open-ended, variable across runs, and difficult to score deterministically. I also explore the limitations of LLM-as-a-judge approaches, including judge disagreement, hallucination, inconsistent standards between judges, and the resulting need to evaluate the evaluators themselves.
Finally, I will consider Pipeline V2: a production system composed of multiple machine-learning models, LLM-powered components, and agentic workflows. I will discuss why component-level accuracy cannot be averaged into system-level reliability, how failures propagate across stages, and how end-to-end evaluation, observability, latency, cost, recovery, and customer outcomes must be combined to determine whether the overall system is successful.
The central lesson is that evaluation must grow with the system. As the unit of intelligence expands from a model to an AI application to a system of systems, benchmarks too must evolve to assess the distribution of real-world behaviour.
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
This talk presents a practical framework for evaluating increasingly complex, non-deterministic AI systems, moving from individual models to complete production pipelines. I begin with a page-stream segmentation model, where standard metrics concealed important failure modes at Wisedocs. For example, we discovered that the accuracy of identifying a document boundary is not the same as the ability to reconstruct the complete document span. Furthemore, errors on some document types carried greater downstream risk than others, and small aggregate improvements did not necessarily translate into meaningful business outcomes. This case study shows how evaluation for us evolved towards task-specific metrics, sensitive failure slices, strict regression testing, and more business-facing metrics of reliability.
I then move to LLM-based systems using Medical Long-Context Reasoning, or MLCR, as a case study. This was a benchmark we developed at Wisedocs earlier this summer. I will examine how to evaluate systems such as agentic chatbots and AI insight-generation services when outputs are open-ended, variable across runs, and difficult to score deterministically. I also explore the limitations of LLM-as-a-judge approaches, including judge disagreement, hallucination, inconsistent standards between judges, and the resulting need to evaluate the evaluators themselves.
Finally, I will consider Pipeline V2: a production system composed of multiple machine-learning models, LLM-powered components, and agentic workflows. I will discuss why component-level accuracy cannot be averaged into system-level reliability, how failures propagate across stages, and how end-to-end evaluation, observability, latency, cost, recovery, and customer outcomes must be combined to determine whether the overall system is successful.
The central lesson is that evaluation must grow with the system. As the unit of intelligence expands from a model to an AI application to a system of systems, benchmarks too must evolve to assess the distribution of real-world behaviour.
Jake Kang
Jake Kang
ABOUT THE SPEAKER:
“This talk shares lessons from building a self-recovering browser agent for regulated healthcare workflows, where we moved from a single-agent prototype to a multi-step system with optimized inputs, memory, evals, sandboxes, and deterministic code where possible. I’ll walk through how we used production failures to build a faster evaluation and prompt optimization loop, and when to move beyond harness engineering to increasing model capability via post-training.
What You’ll Learn
Attendees will learn components of the agent development lifecycle and how to build an effective harness, and when to fine-tune when you hit a plateau with foundation models:
– how to turn production traces into eval datasets
– how to manage context and memory
– how to build sandboxes that speed up testing and iteration.
“
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
“This talk shares lessons from building a self-recovering browser agent for regulated healthcare workflows, where we moved from a single-agent prototype to a multi-step system with optimized inputs, memory, evals, sandboxes, and deterministic code where possible. I’ll walk through how we used production failures to build a faster evaluation and prompt optimization loop, and when to move beyond harness engineering to increasing model capability via post-training.
What You’ll Learn
Attendees will learn components of the agent development lifecycle and how to build an effective harness, and when to fine-tune when you hit a plateau with foundation models:
– how to turn production traces into eval datasets
– how to manage context and memory
– how to build sandboxes that speed up testing and iteration.
“
Jazmia Henry
Jazmia Henry
ABOUT THE SPEAKER:
There’s a gap between researcher-crafted evaluation frameworks that capture model performance and benchmarks versus how end users actually use AI products. This gap is being exploited in ways that render traditional reward models useless.
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
There’s a gap between researcher-crafted evaluation frameworks that capture model performance and benchmarks versus how end users actually use AI products. This gap is being exploited in ways that render traditional reward models useless.
WHAT YOU’LL LEARN:
First, if your domain has computable ground truth, you do not need human annotators to build a reward signal. Second, scalar reward is a compression that loses information. Decomposing reward into multiple verifiable dimensions exposes failure modes that single-number metrics hide. Third, the exploitation gap is a distribution problem, not a labeling problem. Agents exploit the distance between training distribution and deployment reality, and deterministic verifiers narrow that gap directly.
PREREQUISITE KNOWLEDGE
Siddharth Jain
Siddharth Jain
ABOUT THE SPEAKER:
Teams often give an agent a service credential, add a human approval step, and call the workflow governed. That design breaks down when the agent can revise payloads, retry writes, chain tools, or act across systems with different permission models. The result is an accountability gap: the organization can see that a service account acted, but not necessarily who authorized the business intent, which payload was approved, whether a retry duplicated work, or what changed between proposal and execution.
This session presents a production control model for agent workflows that make consequential writes. It shows how to separate the business-intent identity from the concrete operation; classify tools by impact; issue short-lived, least-privilege credentials only after validation; bind human approval to a canonical payload hash and policy version; execute through controlled services with idempotency keys; and reconcile external state before declaring success. It also covers the evidence record needed to answer four operational questions: who requested the action, what the agent proposed, what policy and human approved, and what the downstream system actually did.
The talk focuses on failure modes that appear after the demo works: stale approvals, overbroad agent permissions, duplicate writes after retries, silent policy changes, and actions whose outcome is unknown. Attendees will leave with a lifecycle they can map to their own platform, concrete interfaces between model output and deterministic controls, and an audit schema that supports incident response, compliance review, and day-to-day operations without turning every agent into a bespoke security project.
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
Teams often give an agent a service credential, add a human approval step, and call the workflow governed. That design breaks down when the agent can revise payloads, retry writes, chain tools, or act across systems with different permission models. The result is an accountability gap: the organization can see that a service account acted, but not necessarily who authorized the business intent, which payload was approved, whether a retry duplicated work, or what changed between proposal and execution.
This session presents a production control model for agent workflows that make consequential writes. It shows how to separate the business-intent identity from the concrete operation; classify tools by impact; issue short-lived, least-privilege credentials only after validation; bind human approval to a canonical payload hash and policy version; execute through controlled services with idempotency keys; and reconcile external state before declaring success. It also covers the evidence record needed to answer four operational questions: who requested the action, what the agent proposed, what policy and human approved, and what the downstream system actually did.
The talk focuses on failure modes that appear after the demo works: stale approvals, overbroad agent permissions, duplicate writes after retries, silent policy changes, and actions whose outcome is unknown. Attendees will leave with a lifecycle they can map to their own platform, concrete interfaces between model output and deterministic controls, and an audit schema that supports incident response, compliance review, and day-to-day operations without turning every agent into a bespoke security project.
WHAT YOU’LL LEARN:
- Do not let the model hold broad standing credentials; mint scoped identity for each authorized operation.
- Bind approval to material fields, a canonical payload hash, and the policy version so edits invalidate stale consent.
- Treat every consequential write as a durable state machine with explicit recovery paths, not as a single tool call.
- Separate retry from reconciliation so an ambiguous timeout does not become a duplicate action.
- Design the audit record as an operating interface: it should reconstruct who requested, what was proposed, what was approved, what executed, and what the downstream system recorded.
PREREQUISITE KNOWLEDGE
Manikandan Paramasivan
Manikandan Paramasivan
ABOUT THE SPEAKER:
Building ML models is not the hard part for most teams anymore. Getting them from a data scientist’s notebook to a governed, monitored, production system and keeping them there without a platform team standing between every developer and every deploy is. It gets harder inside a regulated fintech, where every model decision needs an audit trail, every alert needs a documented escalation path, and every “quick fix” has to survive a compliance review.
This talk walks through how our ML platform team at KOHO, a Canadian fintech, built a developer-first MLOps stack on AWS that treats governance as a byproduct of good engineering rather than a tax on top of it and how we’re now layering AI-assisted workflows (Claude Code skills, agentic guardrails) on top to remove the remaining manual toil from model onboarding, monitoring triage, and incident response.
We’ll cover the architecture, the AI-assisted developer tooling built on top of it, and what changed operationally as the company moved toward a full banking license, where “move fast” and “pass the audit” had to become the same sentence, not a trade-off.
Audience takeaway: a concrete pattern for developer productivity that survives regulatory scrutiny, plus where AI agents genuinely help.
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
Building ML models is not the hard part for most teams anymore. Getting them from a data scientist’s notebook to a governed, monitored, production system and keeping them there without a platform team standing between every developer and every deploy is. It gets harder inside a regulated fintech, where every model decision needs an audit trail, every alert needs a documented escalation path, and every “quick fix” has to survive a compliance review.
This talk walks through how our ML platform team at KOHO, a Canadian fintech, built a developer-first MLOps stack on AWS that treats governance as a byproduct of good engineering rather than a tax on top of it and how we’re now layering AI-assisted workflows (Claude Code skills, agentic guardrails) on top to remove the remaining manual toil from model onboarding, monitoring triage, and incident response.
We’ll cover the architecture, the AI-assisted developer tooling built on top of it, and what changed operationally as the company moved toward a full banking license, where “move fast” and “pass the audit” had to become the same sentence, not a trade-off.
Audience takeaway: a concrete pattern for developer productivity that survives regulatory scrutiny, plus where AI agents genuinely help.
Christopher G. Potts
Christopher G. Potts
ABOUT THE SPEAKER:
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
Upal Saha
Upal Saha
ABOUT THE SPEAKER:
This is a build-and-break lab for engineers who already know that a demo is not a system. In the first fifteen minutes every attendee stands up a working extraction pipeline against a real invoice, with no schema authoring, and gets structured JSON back. Then we spend an hour breaking it the way production does, and fixing each break with a pattern that transfers to any stack.
Break one: the wrong document. We feed a bill of lading glued to an invoice into the invoice pipeline and watch it confidently produce garbage. Fix: classify before you extract, and make the graph deterministic even though every step inside it is a model. Break two: the answer that is probably right. Models do not tell you how confident they are, so we compute per-field confidence, find the fields that fall below 95%, and route those, and only those, to a human whose correction feeds back into the system. Break three: the messy string. “10 cases organic gala apples, 88 ct” has to become one SKU; we show why canonicalization is its own step, how to score matches, and where to set the threshold for review. Break four: the hostile input. Everyone runs an image whose pixels read “ignore all previous instructions” and we discuss, with the result on screen, what it means to treat inbound data as data rather than control. We close by labeling a handful of outputs and running a regression test between two versions of the pipeline, because evaluation is a loop, not a phase.
Attendees leave with a running pipeline in their own account and five patterns they can apply on Monday regardless of vendor: classify-then-extract, confidence scoring for models that lack it, threshold-based exception routing with feedback, semantic security checks on inbound data, and versioned regression testing. The platform used in the room is ours; the lessons are not.
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
This is a build-and-break lab for engineers who already know that a demo is not a system. In the first fifteen minutes every attendee stands up a working extraction pipeline against a real invoice, with no schema authoring, and gets structured JSON back. Then we spend an hour breaking it the way production does, and fixing each break with a pattern that transfers to any stack.
Break one: the wrong document. We feed a bill of lading glued to an invoice into the invoice pipeline and watch it confidently produce garbage. Fix: classify before you extract, and make the graph deterministic even though every step inside it is a model. Break two: the answer that is probably right. Models do not tell you how confident they are, so we compute per-field confidence, find the fields that fall below 95%, and route those, and only those, to a human whose correction feeds back into the system. Break three: the messy string. “10 cases organic gala apples, 88 ct” has to become one SKU; we show why canonicalization is its own step, how to score matches, and where to set the threshold for review. Break four: the hostile input. Everyone runs an image whose pixels read “ignore all previous instructions” and we discuss, with the result on screen, what it means to treat inbound data as data rather than control. We close by labeling a handful of outputs and running a regression test between two versions of the pipeline, because evaluation is a loop, not a phase.
Attendees leave with a running pipeline in their own account and five patterns they can apply on Monday regardless of vendor: classify-then-extract, confidence scoring for models that lack it, threshold-based exception routing with feedback, semantic security checks on inbound data, and versioned regression testing. The platform used in the room is ours; the lessons are not.
Antonio Bustamante
Antonio Bustamante
ABOUT THE SPEAKER:
Everyone building on frontier models hits the same wall: 80% of the way there in a weekend, then an exponentially expensive climb toward the 99%+ that operational systems need. This talk is the honest map of that climb, from a team that now runs AI over millions of documents, images and videos a month for customers in logistics, fleet management, automotive and financial services who need the answer to be right every time.
We will walk through the ladder nobody budgets for: retries, then queues when the model is down for three hours, then rate limits, then discovering that 95% is not enough for transactional data, then discovering that the model cannot tell you how confident it is. We will show two failures from our own production history, a customer whose single engineer racked up $30K of usage in a month because nothing was watching, and our first churn, on 500-page reports with a hundred rows per page, a problem we still consider unsolved. And we will show what we changed: a harness that treats AI as a deterministic step inside durable workflows rather than as an open-ended agent, decisions expressed as a verified tree the model must traverse, algorithmic confidence scoring on top of models that provide none, routing anything under 95% to a human whose verdict feeds back into the system, semantic checks against instructions smuggled into the data, and, counterintuitively, encouraging customers to build their own independent monitoring of us. One customer’s users went from eight to nine hours a week on a task to about thirty minutes, and that number was measured by them, not by us.
Attendees will leave with five patterns they can apply Monday: confidence scoring for models that lack it, decision trees over open-ended prompts, exception routing with feedback loops, semantic security checks on inbound data, and customer-owned evaluation. They will also leave with a thesis we did not start with: chat is single-player AI; the next decade of software is ambient AI that runs the same process a hundred thousand times a day, unattended, and behaves the same way every time.
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
Everyone building on frontier models hits the same wall: 80% of the way there in a weekend, then an exponentially expensive climb toward the 99%+ that operational systems need. This talk is the honest map of that climb, from a team that now runs AI over millions of documents, images and videos a month for customers in logistics, fleet management, automotive and financial services who need the answer to be right every time.
We will walk through the ladder nobody budgets for: retries, then queues when the model is down for three hours, then rate limits, then discovering that 95% is not enough for transactional data, then discovering that the model cannot tell you how confident it is. We will show two failures from our own production history, a customer whose single engineer racked up $30K of usage in a month because nothing was watching, and our first churn, on 500-page reports with a hundred rows per page, a problem we still consider unsolved. And we will show what we changed: a harness that treats AI as a deterministic step inside durable workflows rather than as an open-ended agent, decisions expressed as a verified tree the model must traverse, algorithmic confidence scoring on top of models that provide none, routing anything under 95% to a human whose verdict feeds back into the system, semantic checks against instructions smuggled into the data, and, counterintuitively, encouraging customers to build their own independent monitoring of us. One customer’s users went from eight to nine hours a week on a task to about thirty minutes, and that number was measured by them, not by us.
Attendees will leave with five patterns they can apply Monday: confidence scoring for models that lack it, decision trees over open-ended prompts, exception routing with feedback loops, semantic security checks on inbound data, and customer-owned evaluation. They will also leave with a thesis we did not start with: chat is single-player AI; the next decade of software is ambient AI that runs the same process a hundred thousand times a day, unattended, and behaves the same way every time.
Matt Mazzarell
Matt Mazzarell
ABOUT THE SPEAKER:
“One of the most difficult problems every company faces is understanding its customers completely. Customer lifetime value, attrition risk, and purchase propensity are all solvable with AI/ML — but how do we combine these modeling scores to initiate the right action with the right customer at any point in time?
Agentic applications help us make the best possible decisions when interpreting complex, high-volume signals from our customers. An agentic application gives end users visuals that explain key insights, with an agent in the loop to ensure nothing is missed. Context is everything: when done correctly, the agent always has the appropriate understanding to build an action plan that improves customer health and profitability.
In this session, we’ll show you how to build agentic apps from ideation to a finished product that interacts with customers. You’ll take away practical tips for using agentic coding frameworks, curating complete customer data products, and building customer-facing agents — capped off with a live demo of Teradata’s Customer Lifetime Value Agentic App.”
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
“One of the most difficult problems every company faces is understanding its customers completely. Customer lifetime value, attrition risk, and purchase propensity are all solvable with AI/ML — but how do we combine these modeling scores to initiate the right action with the right customer at any point in time?
Agentic applications help us make the best possible decisions when interpreting complex, high-volume signals from our customers. An agentic application gives end users visuals that explain key insights, with an agent in the loop to ensure nothing is missed. Context is everything: when done correctly, the agent always has the appropriate understanding to build an action plan that improves customer health and profitability.
In this session, we’ll show you how to build agentic apps from ideation to a finished product that interacts with customers. You’ll take away practical tips for using agentic coding frameworks, curating complete customer data products, and building customer-facing agents — capped off with a live demo of Teradata’s Customer Lifetime Value Agentic App.”
Vicente Rubén Del Pino Ruiz
Vicente Rubén Del Pino Ruiz
ABOUT THE SPEAKER:
Single-shot evaluation, the shape every current AI eval framework ships, is structurally blind to the failures that take agents down in production: patience that runs out at turn six, users who abandon silently, tool calls that lose conversation state, partial successes that masquerade as wins. This talk argues for a different shape, drawn from the digital twin literature and how the approach is already applied in aerospace, autonomous vehicles, and civil engineering: simulate the population of users your agent will meet, run them against the agent, watch what breaks before any human sees it.
Three take aways:
- Why current AI agent evaluation cannot detect the failures that hit production. The structural reason prompt-and-grade testing (the shape every current framework ships) is blind to patience, abandonment, conversation state, and partial success.
- What a realistic synthetic user population looks like, drawn from the digital twin literature in other engineering industries. Four properties: continuous trait clusters instead of fixed persona archetypes; probabilistic dropout (real users walk away silently rather than say goodbye); three-way goal scoring (each conversation is scored “achieved,” “partially achieved,” or “not achieved,” instead of just pass/fail); and variable conversation length (each conversation ends when the persona’s patience runs out, not at a fixed turn count).
- How to wire population-scale evaluation into a release gate. Why run-over-run delta on a held-constant population is the regression-detection primitive AI agents actually need, and how to treat it as a governance lever rather than a research experiment.”
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
Single-shot evaluation, the shape every current AI eval framework ships, is structurally blind to the failures that take agents down in production: patience that runs out at turn six, users who abandon silently, tool calls that lose conversation state, partial successes that masquerade as wins. This talk argues for a different shape, drawn from the digital twin literature and how the approach is already applied in aerospace, autonomous vehicles, and civil engineering: simulate the population of users your agent will meet, run them against the agent, watch what breaks before any human sees it.
Three take aways:
- Why current AI agent evaluation cannot detect the failures that hit production. The structural reason prompt-and-grade testing (the shape every current framework ships) is blind to patience, abandonment, conversation state, and partial success.
- What a realistic synthetic user population looks like, drawn from the digital twin literature in other engineering industries. Four properties: continuous trait clusters instead of fixed persona archetypes; probabilistic dropout (real users walk away silently rather than say goodbye); three-way goal scoring (each conversation is scored “achieved,” “partially achieved,” or “not achieved,” instead of just pass/fail); and variable conversation length (each conversation ends when the persona’s patience runs out, not at a fixed turn count).
- How to wire population-scale evaluation into a release gate. Why run-over-run delta on a held-constant population is the regression-detection primitive AI agents actually need, and how to treat it as a governance lever rather than a research experiment.”
Yegor Denisov-Blanch
Yegor Denisov-Blanch
ABOUT THE SPEAKER:
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
Andy McMahon
Andy McMahon
ABOUT THE SPEAKER:
As organisations move from single models to connected agent systems, the challenge shifts to how agents interact, how they are evaluated, and how their behaviour can be monitored in real production environments.
This presentation explores what it takes to operationalise agentic AI in practice – from observability and evaluation to deploying agents safely and reliably across enterprise workflows.
Through real-world examples, we look at how teams are building responsible AgentOps frameworks to scale multi-agent systems in production while maintaining control, trust and measurable business impact.
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
As organisations move from single models to connected agent systems, the challenge shifts to how agents interact, how they are evaluated, and how their behaviour can be monitored in real production environments.
This presentation explores what it takes to operationalise agentic AI in practice – from observability and evaluation to deploying agents safely and reliably across enterprise workflows.
Through real-world examples, we look at how teams are building responsible AgentOps frameworks to scale multi-agent systems in production while maintaining control, trust and measurable business impact.
Deji Andrew
Deji Andrew
ABOUT THE SPEAKER:
A LightGBM model was deployed across production lines to predict and correct overruns mid-job, reducing average overproduction from +24 cases to +2 on historical data. In pilot, the corrections worked: overruns vanished. But a closer look at the production data showed something different. When the order target was reduced by 21 cases, the line overproduced the new target by 18. The system was not being corrected – it was negotiating with the correction, regenerating overrun relative to whatever target the machine was now working toward. Compounding this: corrected jobs left no trace in the training data that a human had intervened, meaning the model, on retraining, would learn to stop flagging the problem it had been built to solve. This talk walks through the pilot data, explains why conservatism in the loss function wasn’t enough, and argues that the real gap isn’t in the model – it’s in the experiment infrastructure that doesn’t exist between a prediction and a PLC. Attendees will leave with a concrete mental model of reflexivity and contamination in industrial ML deployment, a framework for when not to close the automation loop, and specific reasons the standard MLOps playbook (holdouts, impression logging, A/B frameworks) doesn’t transfer to systems that act on physical state.
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
A LightGBM model was deployed across production lines to predict and correct overruns mid-job, reducing average overproduction from +24 cases to +2 on historical data. In pilot, the corrections worked: overruns vanished. But a closer look at the production data showed something different. When the order target was reduced by 21 cases, the line overproduced the new target by 18. The system was not being corrected – it was negotiating with the correction, regenerating overrun relative to whatever target the machine was now working toward. Compounding this: corrected jobs left no trace in the training data that a human had intervened, meaning the model, on retraining, would learn to stop flagging the problem it had been built to solve. This talk walks through the pilot data, explains why conservatism in the loss function wasn’t enough, and argues that the real gap isn’t in the model – it’s in the experiment infrastructure that doesn’t exist between a prediction and a PLC. Attendees will leave with a concrete mental model of reflexivity and contamination in industrial ML deployment, a framework for when not to close the automation loop, and specific reasons the standard MLOps playbook (holdouts, impression logging, A/B frameworks) doesn’t transfer to systems that act on physical state.
Arun Malik
Arun Malik
ABOUT THE SPEAKER:
We redesigned frontline incident response around AI agents that do not just suggest fixes but execute them, across more than 12 million devices and tens of thousands of incidents a month. Structurally, three things changed. First, agents stopped getting broad human-equivalent access and instead call tools through a governed interface, with permissions scoped down to individual functions and parameter values, so a compromised or confused agent has a bounded blast radius. Second, human approval moved from a blanket gate to a targeted one, applied only to sensitive or irreversible actions and routed by how familiar the problem is and how reversible the action is. Third, we stopped paying full model-inference cost for repeated work by promoting an agent’s proven, validated behavior into deterministic playbooks that run at near-zero token cost, which cut agent running cost by more than 70 percent over eight months while incident volume doubled.
Attendees will leave able to decide what an agent should be allowed to do at all, where a human actually adds signal versus just latency, and how to drive the cost of autonomy down over time instead of letting it grow. I will be specific about what broke first and the guardrails we only added after an incident.
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
We redesigned frontline incident response around AI agents that do not just suggest fixes but execute them, across more than 12 million devices and tens of thousands of incidents a month. Structurally, three things changed. First, agents stopped getting broad human-equivalent access and instead call tools through a governed interface, with permissions scoped down to individual functions and parameter values, so a compromised or confused agent has a bounded blast radius. Second, human approval moved from a blanket gate to a targeted one, applied only to sensitive or irreversible actions and routed by how familiar the problem is and how reversible the action is. Third, we stopped paying full model-inference cost for repeated work by promoting an agent’s proven, validated behavior into deterministic playbooks that run at near-zero token cost, which cut agent running cost by more than 70 percent over eight months while incident volume doubled.
Attendees will leave able to decide what an agent should be allowed to do at all, where a human actually adds signal versus just latency, and how to drive the cost of autonomy down over time instead of letting it grow. I will be specific about what broke first and the guardrails we only added after an incident.
Jessica Garson Beauchemin
Jessica Garson Beauchemin
ABOUT THE SPEAKER:
Deploying AI applications often involves containers, GPU configuration, dependency management, and infrastructure scaling. Runpod Flash offers a simpler, code-first approach. It lets you define remote functions and hardware requirements in Python and run them on serverless GPUs.
In this talk, we’ll explore how Flash moves Python workloads from local development to cloud deployment, examine its underlying programming model, and build a GPU-backed endpoint. Along the way, we’ll discuss where Flash fits in the AI development stack, the problems it solves, and the trade-offs developers should consider. Attendees will leave with a practical understanding of how to turn local AI code into a scalable service without needing to become infrastructure experts.
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
Deploying AI applications often involves containers, GPU configuration, dependency management, and infrastructure scaling. Runpod Flash offers a simpler, code-first approach. It lets you define remote functions and hardware requirements in Python and run them on serverless GPUs.
In this talk, we’ll explore how Flash moves Python workloads from local development to cloud deployment, examine its underlying programming model, and build a GPU-backed endpoint. Along the way, we’ll discuss where Flash fits in the AI development stack, the problems it solves, and the trade-offs developers should consider. Attendees will leave with a practical understanding of how to turn local AI code into a scalable service without needing to become infrastructure experts.
Jim Allen Wallace
Jim Allen Wallace
ABOUT THE SPEAKER:
Instacart’s ad-serving feature store served features to real-time inference from a multi-hundred-node managed Valkey deployment split across multiple clusters. The team had built a proxy layer, a querying SDK, and a compact storage format on top of it, and still hit three limits: instability during any cluster mutation, tail latency from read fan-out, and cost that got worse when they split clusters to manage the first two.
The structural change was to stop scaling out and scale up. A multi-threaded engine let the team run much larger instances, so hundreds of nodes became roughly 100, each inference request touched far fewer shards, and average and P99 latency fell 50%. Because the proxy and SDK already hid the datastore from ML engineers, the swap required no application changes and tens of terabytes moved without anyone above the storage layer noticing.
We were the engine vendor, and the first pass did not reach the latency the team expected. Getting there took client-side tuning on their end and engine changes from our engineering team, including to the compactor. Validation ran side by side: identical backfill, traffic shifted from 0% to 100%, gates at each percentile.
Attendees will leave able to test a datastore against the fan-out shape of their own feature store rather than an ops/sec number, trace how engine concurrency sets instance size, which sets fan-out, which sets P99, and run a datastore migration under live traffic with per-percentile gates and a rollback path. They will also hear where the trade-off comes back as the cluster grows.
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
Instacart’s ad-serving feature store served features to real-time inference from a multi-hundred-node managed Valkey deployment split across multiple clusters. The team had built a proxy layer, a querying SDK, and a compact storage format on top of it, and still hit three limits: instability during any cluster mutation, tail latency from read fan-out, and cost that got worse when they split clusters to manage the first two.
The structural change was to stop scaling out and scale up. A multi-threaded engine let the team run much larger instances, so hundreds of nodes became roughly 100, each inference request touched far fewer shards, and average and P99 latency fell 50%. Because the proxy and SDK already hid the datastore from ML engineers, the swap required no application changes and tens of terabytes moved without anyone above the storage layer noticing.
We were the engine vendor, and the first pass did not reach the latency the team expected. Getting there took client-side tuning on their end and engine changes from our engineering team, including to the compactor. Validation ran side by side: identical backfill, traffic shifted from 0% to 100%, gates at each percentile.
Attendees will leave able to test a datastore against the fan-out shape of their own feature store rather than an ops/sec number, trace how engine concurrency sets instance size, which sets fan-out, which sets P99, and run a datastore migration under live traffic with per-percentile gates and a rollback path. They will also hear where the trade-off comes back as the cluster grows.
Zachary Hamilton
Zachary Hamilton
ABOUT THE SPEAKER:
Building agentic software requires more than adding an LLM to the traditional software development lifecycle. When behavior becomes non-deterministic, teams need a new feedback loop connecting what happens in production to how they evaluate, debug, and improve their systems. This talk will show how to build that loop: identifying real failure modes from production traces and human feedback, turning them into reproducible datasets and evaluations, defining success using both technical and business outcomes, and using automated and human judges to measure whether changes actually improve the system. We’ll also explore where automated evaluation breaks down, how to calibrate LLM-as-a-judge, and how continuous evaluation helps teams detect regressions and emerging behavior. Attendees will leave with a practical framework for moving from production > failure discovery > datasets > evals > iteration > production, turning the development of agentic systems into a measurable engineering discipline.
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
Building agentic software requires more than adding an LLM to the traditional software development lifecycle. When behavior becomes non-deterministic, teams need a new feedback loop connecting what happens in production to how they evaluate, debug, and improve their systems. This talk will show how to build that loop: identifying real failure modes from production traces and human feedback, turning them into reproducible datasets and evaluations, defining success using both technical and business outcomes, and using automated and human judges to measure whether changes actually improve the system. We’ll also explore where automated evaluation breaks down, how to calibrate LLM-as-a-judge, and how continuous evaluation helps teams detect regressions and emerging behavior. Attendees will leave with a practical framework for moving from production > failure discovery > datasets > evals > iteration > production, turning the development of agentic systems into a measurable engineering discipline.
Keynote
D. Sculley
D. Sculley
ABOUT THE SPEAKER:
We spend a lot of time thinking about operational issues in AI related to deployment, somewhat less time thinking about adoption, or (dare we say it) acceptance. This talk will touch on some technical pieces in the current AI ops landscape including streaming systems, planning, latency, and on-device models, but most of the time will be spent looking at ways we can move beyond the stale framing of a chatbot, assistant, or customer service agent.
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
We spend a lot of time thinking about operational issues in AI related to deployment, somewhat less time thinking about adoption, or (dare we say it) acceptance. This talk will touch on some technical pieces in the current AI ops landscape including streaming systems, planning, latency, and on-device models, but most of the time will be spent looking at ways we can move beyond the stale framing of a chatbot, assistant, or customer service agent.
Agent Harness Engineering
Arun Malik
Arun Malik
ABOUT THE SPEAKER:
We redesigned frontline incident response around AI agents that do not just suggest fixes but execute them, across more than 12 million devices and tens of thousands of incidents a month. Structurally, three things changed. First, agents stopped getting broad human-equivalent access and instead call tools through a governed interface, with permissions scoped down to individual functions and parameter values, so a compromised or confused agent has a bounded blast radius. Second, human approval moved from a blanket gate to a targeted one, applied only to sensitive or irreversible actions and routed by how familiar the problem is and how reversible the action is. Third, we stopped paying full model-inference cost for repeated work by promoting an agent’s proven, validated behavior into deterministic playbooks that run at near-zero token cost, which cut agent running cost by more than 70 percent over eight months while incident volume doubled.
Attendees will leave able to decide what an agent should be allowed to do at all, where a human actually adds signal versus just latency, and how to drive the cost of autonomy down over time instead of letting it grow. I will be specific about what broke first and the guardrails we only added after an incident.
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
We redesigned frontline incident response around AI agents that do not just suggest fixes but execute them, across more than 12 million devices and tens of thousands of incidents a month. Structurally, three things changed. First, agents stopped getting broad human-equivalent access and instead call tools through a governed interface, with permissions scoped down to individual functions and parameter values, so a compromised or confused agent has a bounded blast radius. Second, human approval moved from a blanket gate to a targeted one, applied only to sensitive or irreversible actions and routed by how familiar the problem is and how reversible the action is. Third, we stopped paying full model-inference cost for repeated work by promoting an agent’s proven, validated behavior into deterministic playbooks that run at near-zero token cost, which cut agent running cost by more than 70 percent over eight months while incident volume doubled.
Attendees will leave able to decide what an agent should be allowed to do at all, where a human actually adds signal versus just latency, and how to drive the cost of autonomy down over time instead of letting it grow. I will be specific about what broke first and the guardrails we only added after an incident.
Jake Kang
Jake Kang
ABOUT THE SPEAKER:
“This talk shares lessons from building a self-recovering browser agent for regulated healthcare workflows, where we moved from a single-agent prototype to a multi-step system with optimized inputs, memory, evals, sandboxes, and deterministic code where possible. I’ll walk through how we used production failures to build a faster evaluation and prompt optimization loop, and when to move beyond harness engineering to increasing model capability via post-training.
What You’ll Learn
Attendees will learn components of the agent development lifecycle and how to build an effective harness, and when to fine-tune when you hit a plateau with foundation models:
– how to turn production traces into eval datasets
– how to manage context and memory
– how to build sandboxes that speed up testing and iteration.
“
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
“This talk shares lessons from building a self-recovering browser agent for regulated healthcare workflows, where we moved from a single-agent prototype to a multi-step system with optimized inputs, memory, evals, sandboxes, and deterministic code where possible. I’ll walk through how we used production failures to build a faster evaluation and prompt optimization loop, and when to move beyond harness engineering to increasing model capability via post-training.
What You’ll Learn
Attendees will learn components of the agent development lifecycle and how to build an effective harness, and when to fine-tune when you hit a plateau with foundation models:
– how to turn production traces into eval datasets
– how to manage context and memory
– how to build sandboxes that speed up testing and iteration.
“
Robert Lewis
Robert Lewis
ABOUT THE SPEAKER:
Large language model (LLM) agents can now draft executable business logic — compliance rules, audit checks, detection heuristics — orders of magnitude faster than human engineers. In regulated domains, however, no such logic ships without a human domain expert’s sign-off. We identify a subtle and, in our experience, pervasive failure mode in this human-in-the-loop pattern: review/serving skew, in which the evidence a domain expert reviews is generated by a re-implementation of the candidate logic (a survey script, a notebook, a prompt-level approximation) rather than by the production engine that will ultimately execute it. When the two implementations diverge — and they reliably do — expert approval validates the wrong behavior. We present a discipline that eliminates this failure mode by construction: candidate rules are drafted directly into the production repository in an inactive-but-runnable state, all review evidence is generated by the production execution path, refinement edits the production artifact itself, and activation is reduced to a one-line data change. We describe the discipline in the context of a deployed multi-agent system that authors, validates, and ships declarative compliance rules for automated audit of collision-repair insurance estimates, evaluated against a corpus of approximately 497,000 production documents. The system couples eight specialized LLM agents with deterministic safety rails — golden-fixture regression, an activation allowlist, single-rule change isolation, and corpus-governance constraints — and drives refinement through a bounded, measurable convergence loop against expert false-positive verdicts. We formalize the discipline as five principles, catalog the anti-patterns it forbids, report operational experience from sixteen production rules, and argue that the pattern transfers directly to clinical decision support, financial transaction monitoring, legal document review, and content moderation.
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
Large language model (LLM) agents can now draft executable business logic — compliance rules, audit checks, detection heuristics — orders of magnitude faster than human engineers. In regulated domains, however, no such logic ships without a human domain expert’s sign-off. We identify a subtle and, in our experience, pervasive failure mode in this human-in-the-loop pattern: review/serving skew, in which the evidence a domain expert reviews is generated by a re-implementation of the candidate logic (a survey script, a notebook, a prompt-level approximation) rather than by the production engine that will ultimately execute it. When the two implementations diverge — and they reliably do — expert approval validates the wrong behavior. We present a discipline that eliminates this failure mode by construction: candidate rules are drafted directly into the production repository in an inactive-but-runnable state, all review evidence is generated by the production execution path, refinement edits the production artifact itself, and activation is reduced to a one-line data change. We describe the discipline in the context of a deployed multi-agent system that authors, validates, and ships declarative compliance rules for automated audit of collision-repair insurance estimates, evaluated against a corpus of approximately 497,000 production documents. The system couples eight specialized LLM agents with deterministic safety rails — golden-fixture regression, an activation allowlist, single-rule change isolation, and corpus-governance constraints — and drives refinement through a bounded, measurable convergence loop against expert false-positive verdicts. We formalize the discipline as five principles, catalog the anti-patterns it forbids, report operational experience from sixteen production rules, and argue that the pattern transfers directly to clinical decision support, financial transaction monitoring, legal document review, and content moderation.
WHAT YOU’LL LEARN:
Robert Joel Lewis is a computational social scientist and senior AI engineer whose work bridges artificial intelligence, statistics, software development, and human behavior. Across a career spanning academia and industry, he has published more than 30 scholarly works and built AI-driven products for ecommerce, entertainment, mobile gaming, and collision repair. His research examines entertainment as an attention-capturing technology, including a forthcoming article in the Journal of Media Psychology. Based in Austin, Texas, he specializes in turning complex, messy problems into rigorous systems that people can actually use.
PREREQUISITE KNOWLEDGE
Amit Kumar Padhy
Amit Kumar Padhy
ABOUT THE SPEAKER:
Modern commerce platforms don’t fail because of missing features, they fail at the seams.
A product is created in Catalog, but pricing is incomplete. Promotions don’t qualify. Tax blocks specific regions. Localization lags. The system says “”launched,” but the business knows it isn’t. These aren’t edge cases, they are the steady state of distributed commerce.
This session replaces traditional workflow orchestration with a multi-agent, swarm-based execution model powered by LLMs and agentic AI, coordinating Pricing, Catalog, Promotions, Tax, and Compliance in real time.
Agent Roles. Planner Agents decompose onboarding goals into executable plans using ReAct-style tool-aware reasoning and function calling. Domain Agents for Pricing, Catalog, and Compliance execute directly against APIs, Pricing Runtime, Offer Systems, Billing Preview, Tax engines. Validator Agents enforce policy rules, regional compliance, and pricing integrity at every step. A Coordinator Agent maintains shared state and resolves cross-domain conflicts through blackboard-pattern memory.
Execution Model. Agents subscribe to domain events, SKU created, price missing, policy violation, via streaming infrastructure. LLM-backed planners generate dynamic execution graphs, not fixed DAGs. Each agent invokes APIs as tools, binding service calls to reasoning steps. Coordination flows through a shared state store and event-driven updates, enabling adaptive planning under partial data, cross-service dependency resolution, and parallel execution with conflict detection.
Failure Handling. No blind retries. LLM-guided compensations roll back partial pricing, recompute discounts, and re-trigger downstream syncs based on causal analysis. No opaque errors. Every decision produces a chain-of-thought trace with tool invocation logs and reasoning outputs. No brittle state. Recovery is idempotent and event-sourced.
Onboarding Flow. Catalog Agents handle onboarding for products by ensuring the product definition is created, and the Pricing agent ensures SKUs are priced. Promotion Agents conduct market checks, provide qualified promotions and their onboarding, and resolve discounts using a rule-based and LLM-hybrid evaluation.
Compliance Agents validate regulatory constraints per geography. Localization Agents assess market-readiness signals, and product merchandising is localized accordingly. The Coordinator drives convergence toward “sellable state”.
Engineering Stack. Kafka-backed event contracts for agent triggers. Function-calling LLM agents with API abstraction layers. Vector DB plus state store for context propagation. Human-in-the-loop decision gates only at high-risk boundaries. Immutable execution logs tracing every action and compensation. Observability is built around decision reasoning, not just service health.
Takeaway. A working blueprint for LLM-powered agent swarms that handle uncertainty across distributed commerce, recover through intelligent compensation, and ensure products ship globally, not just deploy.
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
Modern commerce platforms don’t fail because of missing features, they fail at the seams.
A product is created in Catalog, but pricing is incomplete. Promotions don’t qualify. Tax blocks specific regions. Localization lags. The system says “”launched,” but the business knows it isn’t. These aren’t edge cases, they are the steady state of distributed commerce.
This session replaces traditional workflow orchestration with a multi-agent, swarm-based execution model powered by LLMs and agentic AI, coordinating Pricing, Catalog, Promotions, Tax, and Compliance in real time.
Agent Roles. Planner Agents decompose onboarding goals into executable plans using ReAct-style tool-aware reasoning and function calling. Domain Agents for Pricing, Catalog, and Compliance execute directly against APIs, Pricing Runtime, Offer Systems, Billing Preview, Tax engines. Validator Agents enforce policy rules, regional compliance, and pricing integrity at every step. A Coordinator Agent maintains shared state and resolves cross-domain conflicts through blackboard-pattern memory.
Execution Model. Agents subscribe to domain events, SKU created, price missing, policy violation, via streaming infrastructure. LLM-backed planners generate dynamic execution graphs, not fixed DAGs. Each agent invokes APIs as tools, binding service calls to reasoning steps. Coordination flows through a shared state store and event-driven updates, enabling adaptive planning under partial data, cross-service dependency resolution, and parallel execution with conflict detection.
Failure Handling. No blind retries. LLM-guided compensations roll back partial pricing, recompute discounts, and re-trigger downstream syncs based on causal analysis. No opaque errors. Every decision produces a chain-of-thought trace with tool invocation logs and reasoning outputs. No brittle state. Recovery is idempotent and event-sourced.
Onboarding Flow. Catalog Agents handle onboarding for products by ensuring the product definition is created, and the Pricing agent ensures SKUs are priced. Promotion Agents conduct market checks, provide qualified promotions and their onboarding, and resolve discounts using a rule-based and LLM-hybrid evaluation.
Compliance Agents validate regulatory constraints per geography. Localization Agents assess market-readiness signals, and product merchandising is localized accordingly. The Coordinator drives convergence toward “sellable state”.
Engineering Stack. Kafka-backed event contracts for agent triggers. Function-calling LLM agents with API abstraction layers. Vector DB plus state store for context propagation. Human-in-the-loop decision gates only at high-risk boundaries. Immutable execution logs tracing every action and compensation. Observability is built around decision reasoning, not just service health.
Takeaway. A working blueprint for LLM-powered agent swarms that handle uncertainty across distributed commerce, recover through intelligent compensation, and ensure products ship globally, not just deploy.
Ishaan Sehgal
Ishaan Sehgal
ABOUT THE SPEAKER:
At Omnara, long-running agent sessions exposed a failure mode hidden by short-lived agents: the agent’s state was coupled to the process, model provider, or runtime executing it. When a worker crashed, a client disconnected, or execution moved to another machine, recovering the agent’s work became brittle.
We redesigned agent execution around an append-only session log as the source of truth. Model outputs, tool calls and results, user interventions, and other state transitions are persisted as ordered events; the live agent state becomes a replayable projection of that history.
This structural change decouples agent identity from any one model, worker, or harness. It enables crash recovery, resumability, branching, real-time observability, and migration across models and machines, but introduces difficult systems questions around ordering, idempotency, partial tool execution, checkpointing, compaction, concurrent writers, and log ownership.
This session walks through the production failures that motivated the redesign, the resulting event model and recovery path, and the tradeoffs we encountered. Using production traces and forced-failure examples, we will show what happens when a worker dies mid-tool call, how a session resumes under a different runtime, and how replay costs change as histories grow.
Attendees will leave with a practical blueprint for deciding what belongs in the durable log, what can be recomputed, and how to build agent sessions that survive process, machine, and model failure without locking continuity to a single provider.
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
At Omnara, long-running agent sessions exposed a failure mode hidden by short-lived agents: the agent’s state was coupled to the process, model provider, or runtime executing it. When a worker crashed, a client disconnected, or execution moved to another machine, recovering the agent’s work became brittle.
We redesigned agent execution around an append-only session log as the source of truth. Model outputs, tool calls and results, user interventions, and other state transitions are persisted as ordered events; the live agent state becomes a replayable projection of that history.
This structural change decouples agent identity from any one model, worker, or harness. It enables crash recovery, resumability, branching, real-time observability, and migration across models and machines, but introduces difficult systems questions around ordering, idempotency, partial tool execution, checkpointing, compaction, concurrent writers, and log ownership.
This session walks through the production failures that motivated the redesign, the resulting event model and recovery path, and the tradeoffs we encountered. Using production traces and forced-failure examples, we will show what happens when a worker dies mid-tool call, how a session resumes under a different runtime, and how replay costs change as histories grow.
Attendees will leave with a practical blueprint for deciding what belongs in the durable log, what can be recomputed, and how to build agent sessions that survive process, machine, and model failure without locking continuity to a single provider.
Kumaran Ponnambalam
Kumaran Ponnambalam
ABOUT THE SPEAKER:
This session examines how a traditionally static enterprise agent workflow was redesigned around AI-driven adaptation. In most enterprise settings, agent behavior is defined through fixed prompts, rules, and workflows that apply broadly across users and tenants. We redesigned that model into an adaptive agent architecture built around explicit policy layers, feedback loops, and governed personalization. Structurally, this changed the system from one-time configuration to continuous learning: policies were separated from model reasoning, policy scope was organized across global, tenant, and user levels, and agent behavior was connected to both explicit and implicit feedback signals that could influence future decisions.
The session will focus on what changed in the system architecture and operating model to support this shift. It will cover how policy-driven personalization was introduced, how feedback became part of the runtime and improvement loop, and how adaptation was managed without losing enterprise control. It will also discuss the practical challenges that emerged when moving from static agents to adaptive ones, including handling conflicting user and tenant needs, keeping personalization aligned with core guardrails, and making behavior changes observable and manageable in production.
Attendees will walk away with a practical understanding of how to redesign enterprise agent workflows around adaptive AI capabilities. They will learn the core concepts behind Personalized Experiential Learning, the structural changes needed to support policy-based adaptation, and the key lessons and best practices for making adaptive agents workable in real enterprise environments. They should leave better equipped to think about how to introduce governed personalization, feedback-driven improvement, and policy evolution into their own AI agent systems.
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
This session examines how a traditionally static enterprise agent workflow was redesigned around AI-driven adaptation. In most enterprise settings, agent behavior is defined through fixed prompts, rules, and workflows that apply broadly across users and tenants. We redesigned that model into an adaptive agent architecture built around explicit policy layers, feedback loops, and governed personalization. Structurally, this changed the system from one-time configuration to continuous learning: policies were separated from model reasoning, policy scope was organized across global, tenant, and user levels, and agent behavior was connected to both explicit and implicit feedback signals that could influence future decisions.
The session will focus on what changed in the system architecture and operating model to support this shift. It will cover how policy-driven personalization was introduced, how feedback became part of the runtime and improvement loop, and how adaptation was managed without losing enterprise control. It will also discuss the practical challenges that emerged when moving from static agents to adaptive ones, including handling conflicting user and tenant needs, keeping personalization aligned with core guardrails, and making behavior changes observable and manageable in production.
Attendees will walk away with a practical understanding of how to redesign enterprise agent workflows around adaptive AI capabilities. They will learn the core concepts behind Personalized Experiential Learning, the structural changes needed to support policy-based adaptation, and the key lessons and best practices for making adaptive agents workable in real enterprise environments. They should leave better equipped to think about how to introduce governed personalization, feedback-driven improvement, and policy evolution into their own AI agent systems.
Agent Deployment & Observability
Maitrik Patel
Maitrik Patel
ABOUT THE SPEAKER:
Running ML training, AI inference, and agent workflows on three separate orchestration systems is an operational debt that compounds with every new workload type – one system per abstraction means three on-call rotations, three observability stacks, and three onboarding paths. We collapsed this to a single execution layer by designing a type-safe task interface that accommodates non-deterministic agent execution alongside deterministic ML pipelines, backed by container-native isolation and a workload-aware shared scheduler. The unification exposed a hard failure class that static ML infrastructure never encounters: agent tasks are not idempotent, and ML-derived retry semantics caused downstream state corruption in early production until we rebuilt retry logic around explicit checkpointing and action journals. Attendees leave with a four-layer architectural blueprint for unifying these workload types, an honest account of the scheduling contention and retry failures encountered in production, and the specific interface design decisions that made incremental migration rather than forced rewrite the path to adoption.
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
Running ML training, AI inference, and agent workflows on three separate orchestration systems is an operational debt that compounds with every new workload type – one system per abstraction means three on-call rotations, three observability stacks, and three onboarding paths. We collapsed this to a single execution layer by designing a type-safe task interface that accommodates non-deterministic agent execution alongside deterministic ML pipelines, backed by container-native isolation and a workload-aware shared scheduler. The unification exposed a hard failure class that static ML infrastructure never encounters: agent tasks are not idempotent, and ML-derived retry semantics caused downstream state corruption in early production until we rebuilt retry logic around explicit checkpointing and action journals. Attendees leave with a four-layer architectural blueprint for unifying these workload types, an honest account of the scheduling contention and retry failures encountered in production, and the specific interface design decisions that made incremental migration rather than forced rewrite the path to adoption.
Evals and Benchmarks
Manikandan Paramasivan
Manikandan Paramasivan
ABOUT THE SPEAKER:
Building ML models is not the hard part for most teams anymore. Getting them from a data scientist’s notebook to a governed, monitored, production system and keeping them there without a platform team standing between every developer and every deploy is. It gets harder inside a regulated fintech, where every model decision needs an audit trail, every alert needs a documented escalation path, and every “quick fix” has to survive a compliance review.
This talk walks through how our ML platform team at KOHO, a Canadian fintech, built a developer-first MLOps stack on AWS that treats governance as a byproduct of good engineering rather than a tax on top of it and how we’re now layering AI-assisted workflows (Claude Code skills, agentic guardrails) on top to remove the remaining manual toil from model onboarding, monitoring triage, and incident response.
We’ll cover the architecture, the AI-assisted developer tooling built on top of it, and what changed operationally as the company moved toward a full banking license, where “move fast” and “pass the audit” had to become the same sentence, not a trade-off.
Audience takeaway: a concrete pattern for developer productivity that survives regulatory scrutiny, plus where AI agents genuinely help.
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
Building ML models is not the hard part for most teams anymore. Getting them from a data scientist’s notebook to a governed, monitored, production system and keeping them there without a platform team standing between every developer and every deploy is. It gets harder inside a regulated fintech, where every model decision needs an audit trail, every alert needs a documented escalation path, and every “quick fix” has to survive a compliance review.
This talk walks through how our ML platform team at KOHO, a Canadian fintech, built a developer-first MLOps stack on AWS that treats governance as a byproduct of good engineering rather than a tax on top of it and how we’re now layering AI-assisted workflows (Claude Code skills, agentic guardrails) on top to remove the remaining manual toil from model onboarding, monitoring triage, and incident response.
We’ll cover the architecture, the AI-assisted developer tooling built on top of it, and what changed operationally as the company moved toward a full banking license, where “move fast” and “pass the audit” had to become the same sentence, not a trade-off.
Audience takeaway: a concrete pattern for developer productivity that survives regulatory scrutiny, plus where AI agents genuinely help.
Aryan Dhar
Aryan Dhar
ABOUT THE SPEAKER:
This talk presents a practical framework for evaluating increasingly complex, non-deterministic AI systems, moving from individual models to complete production pipelines. I begin with a page-stream segmentation model, where standard metrics concealed important failure modes at Wisedocs. For example, we discovered that the accuracy of identifying a document boundary is not the same as the ability to reconstruct the complete document span. Furthemore, errors on some document types carried greater downstream risk than others, and small aggregate improvements did not necessarily translate into meaningful business outcomes. This case study shows how evaluation for us evolved towards task-specific metrics, sensitive failure slices, strict regression testing, and more business-facing metrics of reliability.
I then move to LLM-based systems using Medical Long-Context Reasoning, or MLCR, as a case study. This was a benchmark we developed at Wisedocs earlier this summer. I will examine how to evaluate systems such as agentic chatbots and AI insight-generation services when outputs are open-ended, variable across runs, and difficult to score deterministically. I also explore the limitations of LLM-as-a-judge approaches, including judge disagreement, hallucination, inconsistent standards between judges, and the resulting need to evaluate the evaluators themselves.
Finally, I will consider Pipeline V2: a production system composed of multiple machine-learning models, LLM-powered components, and agentic workflows. I will discuss why component-level accuracy cannot be averaged into system-level reliability, how failures propagate across stages, and how end-to-end evaluation, observability, latency, cost, recovery, and customer outcomes must be combined to determine whether the overall system is successful.
The central lesson is that evaluation must grow with the system. As the unit of intelligence expands from a model to an AI application to a system of systems, benchmarks too must evolve to assess the distribution of real-world behaviour.
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
This talk presents a practical framework for evaluating increasingly complex, non-deterministic AI systems, moving from individual models to complete production pipelines. I begin with a page-stream segmentation model, where standard metrics concealed important failure modes at Wisedocs. For example, we discovered that the accuracy of identifying a document boundary is not the same as the ability to reconstruct the complete document span. Furthemore, errors on some document types carried greater downstream risk than others, and small aggregate improvements did not necessarily translate into meaningful business outcomes. This case study shows how evaluation for us evolved towards task-specific metrics, sensitive failure slices, strict regression testing, and more business-facing metrics of reliability.
I then move to LLM-based systems using Medical Long-Context Reasoning, or MLCR, as a case study. This was a benchmark we developed at Wisedocs earlier this summer. I will examine how to evaluate systems such as agentic chatbots and AI insight-generation services when outputs are open-ended, variable across runs, and difficult to score deterministically. I also explore the limitations of LLM-as-a-judge approaches, including judge disagreement, hallucination, inconsistent standards between judges, and the resulting need to evaluate the evaluators themselves.
Finally, I will consider Pipeline V2: a production system composed of multiple machine-learning models, LLM-powered components, and agentic workflows. I will discuss why component-level accuracy cannot be averaged into system-level reliability, how failures propagate across stages, and how end-to-end evaluation, observability, latency, cost, recovery, and customer outcomes must be combined to determine whether the overall system is successful.
The central lesson is that evaluation must grow with the system. As the unit of intelligence expands from a model to an AI application to a system of systems, benchmarks too must evolve to assess the distribution of real-world behaviour.
Drew Crawford
Drew Crawford
ABOUT THE SPEAKER:
Somewhere in a datacenter right now, one frontier model is telling another that it’s the town doctor. It is not the town doctor. It killed someone four turns ago, and it’s about to get a third model lynched for it.
Most LLM benchmarks are exams: a model alone in a room with a test paper. Mafia is a room where some of the agents are lying, everyone knows some of them are lying, and the game is figuring out which. It demands recursive theory of mind, deception that stays consistent under adversarial re-reading, and long-context discipline — and it punishes output indiscipline like no static eval can: a model that rambles or breaks format doesn’t get a bad grade, it gets voted out. The other players are the evaluation harness.
This talk covers what it took to run AI Bot Mafia in production: 40,000 lines of Rust; the streaming pathologies of 104 model configurations (chunks that split UTF-8 characters mid-byte, keepalive-only hangs no single timeout catches, malformed tool calls that poison history); and cost governance where prompt-cache economics made cache locality a game-design constraint, enforced by spend fuses and a concurrency queue.
Then the leaderboard: why win rates and Elo are statistically indefensible here, and what works instead — ridge logistic regression over per-seat signed features, bootstrapped 200×, ranked by the p10 lower bound, with a machine-readable reliability state that says “insufficient data” out loud.
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
Somewhere in a datacenter right now, one frontier model is telling another that it’s the town doctor. It is not the town doctor. It killed someone four turns ago, and it’s about to get a third model lynched for it.
Most LLM benchmarks are exams: a model alone in a room with a test paper. Mafia is a room where some of the agents are lying, everyone knows some of them are lying, and the game is figuring out which. It demands recursive theory of mind, deception that stays consistent under adversarial re-reading, and long-context discipline — and it punishes output indiscipline like no static eval can: a model that rambles or breaks format doesn’t get a bad grade, it gets voted out. The other players are the evaluation harness.
This talk covers what it took to run AI Bot Mafia in production: 40,000 lines of Rust; the streaming pathologies of 104 model configurations (chunks that split UTF-8 characters mid-byte, keepalive-only hangs no single timeout catches, malformed tool calls that poison history); and cost governance where prompt-cache economics made cache locality a game-design constraint, enforced by spend fuses and a concurrency queue.
Then the leaderboard: why win rates and Elo are statistically indefensible here, and what works instead — ridge logistic regression over per-seat signed features, bootstrapped 200×, ranked by the p10 lower bound, with a machine-readable reliability state that says “insufficient data” out loud.
Evaluation & Testing of Non-Deterministic Systems
Deji Andrew
Deji Andrew
ABOUT THE SPEAKER:
A LightGBM model was deployed across production lines to predict and correct overruns mid-job, reducing average overproduction from +24 cases to +2 on historical data. In pilot, the corrections worked: overruns vanished. But a closer look at the production data showed something different. When the order target was reduced by 21 cases, the line overproduced the new target by 18. The system was not being corrected – it was negotiating with the correction, regenerating overrun relative to whatever target the machine was now working toward. Compounding this: corrected jobs left no trace in the training data that a human had intervened, meaning the model, on retraining, would learn to stop flagging the problem it had been built to solve. This talk walks through the pilot data, explains why conservatism in the loss function wasn’t enough, and argues that the real gap isn’t in the model – it’s in the experiment infrastructure that doesn’t exist between a prediction and a PLC. Attendees will leave with a concrete mental model of reflexivity and contamination in industrial ML deployment, a framework for when not to close the automation loop, and specific reasons the standard MLOps playbook (holdouts, impression logging, A/B frameworks) doesn’t transfer to systems that act on physical state.
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
A LightGBM model was deployed across production lines to predict and correct overruns mid-job, reducing average overproduction from +24 cases to +2 on historical data. In pilot, the corrections worked: overruns vanished. But a closer look at the production data showed something different. When the order target was reduced by 21 cases, the line overproduced the new target by 18. The system was not being corrected – it was negotiating with the correction, regenerating overrun relative to whatever target the machine was now working toward. Compounding this: corrected jobs left no trace in the training data that a human had intervened, meaning the model, on retraining, would learn to stop flagging the problem it had been built to solve. This talk walks through the pilot data, explains why conservatism in the loss function wasn’t enough, and argues that the real gap isn’t in the model – it’s in the experiment infrastructure that doesn’t exist between a prediction and a PLC. Attendees will leave with a concrete mental model of reflexivity and contamination in industrial ML deployment, a framework for when not to close the automation loop, and specific reasons the standard MLOps playbook (holdouts, impression logging, A/B frameworks) doesn’t transfer to systems that act on physical state.
Cost Management and ROI
Nadia Rauch
Nadia Rauch
ABOUT THE SPEAKER:
We built a multi-agent pipeline to automate a complex, document-intensive enterprise workflow. The initial system ran sequentially through a chain of specialized agent roles, each handling a distinct task in the process, using a single frontier model throughout. It worked. It was also slow and expensive, and we didn’t know why.
Rather than optimize blindly, we profiled first. The results were not where we expected: 67% of total latency came from a small minority of the agent roles, and the primary bottleneck was not the model — it was sequential chaining: independent work was being processed one step at a time instead of concurrently. Identifying and parallelizing the roles with no inter-dependency reduced end-to-end runtime by roughly 40%.*
The second experiment compared model tiers (frontier, mid-tier, lightweight) across each agent role independently, measuring accuracy, latency, and cost per role. The finding cuts against the default assumption: the roles that appeared most cognitively demanding required frontier models, but the accuracy gap was smaller than expected. The roles where model downgrade failed were the ones responsible for precise structured extraction — tasks where errors propagate silently downstream and surface only at the output. Swapping to mid-tier models on roles that tolerated it reduced per-run cost by ~40%, with an estimated <2% accuracy impact at the role level and <1% at the pipeline output level.*
The third experiment compared orchestration harnesses — evaluating how the choice of agentic framework affects runtime overhead, observability, and the ease of implementing the parallelization and model-swap changes described above. Framework choice turned out to matter more than expected for operational concerns: debugging multi-agent failures, tracing costs per agent, and modifying execution flow without rewriting pipeline logic.
Operating in a regulated enterprise environment added one constraint worth naming: every agent decision needs to be auditable. This ruled out certain optimization shortcuts that would have been acceptable in other contexts and shaped how we defined “accurate enough” per role.
This talk covers the profiling methodology, the per-role model comparison framework, the parallelization decisions, the harness comparison, and the operational lessons — including what we’d instrument from day one if we rebuilt the system today.
*Cost and accuracy figures for the parallelization and model-tier experiments are preliminary; the formal evaluation is in progress and will be updated with measured results before the talk.
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
We built a multi-agent pipeline to automate a complex, document-intensive enterprise workflow. The initial system ran sequentially through a chain of specialized agent roles, each handling a distinct task in the process, using a single frontier model throughout. It worked. It was also slow and expensive, and we didn’t know why.
Rather than optimize blindly, we profiled first. The results were not where we expected: 67% of total latency came from a small minority of the agent roles, and the primary bottleneck was not the model — it was sequential chaining: independent work was being processed one step at a time instead of concurrently. Identifying and parallelizing the roles with no inter-dependency reduced end-to-end runtime by roughly 40%.*
The second experiment compared model tiers (frontier, mid-tier, lightweight) across each agent role independently, measuring accuracy, latency, and cost per role. The finding cuts against the default assumption: the roles that appeared most cognitively demanding required frontier models, but the accuracy gap was smaller than expected. The roles where model downgrade failed were the ones responsible for precise structured extraction — tasks where errors propagate silently downstream and surface only at the output. Swapping to mid-tier models on roles that tolerated it reduced per-run cost by ~40%, with an estimated <2% accuracy impact at the role level and <1% at the pipeline output level.*
The third experiment compared orchestration harnesses — evaluating how the choice of agentic framework affects runtime overhead, observability, and the ease of implementing the parallelization and model-swap changes described above. Framework choice turned out to matter more than expected for operational concerns: debugging multi-agent failures, tracing costs per agent, and modifying execution flow without rewriting pipeline logic.
Operating in a regulated enterprise environment added one constraint worth naming: every agent decision needs to be auditable. This ruled out certain optimization shortcuts that would have been acceptable in other contexts and shaped how we defined “accurate enough” per role.
This talk covers the profiling methodology, the per-role model comparison framework, the parallelization decisions, the harness comparison, and the operational lessons — including what we’d instrument from day one if we rebuilt the system today.
*Cost and accuracy figures for the parallelization and model-tier experiments are preliminary; the formal evaluation is in progress and will be updated with measured results before the talk.
Tony Blank
Tony Blank
ABOUT THE SPEAKER:
We stopped equipping humans and started automating their worst hour. Instead of teaching people to fish, we built the pond: precise background agents pointed at repeated, skill-independent time-sinks where ROI was measurable on day one. The first build was a bug-triage agent wired to GitHub, Jira, and Notion; every agent that followed was measured the same way — Time Saved Per Task (TSPT) per week — against an amortized cost model (build + inference + maintenance vs. hours saved × loaded hourly cost).
Structurally, the change was moving from a portfolio of per-person tool subscriptions to a portfolio of scoped, measurable background agents with a pre-build attribution conversation baked in. The decision rubric — Sharpness × Cadence × Skill Tax → Pond / Boat / Net / Wade — tells you when to automate fully, when to augment, when to keep a human in the loop, and when the math doesn’t work yet.
Attendees will leave able to: score their own candidate workflows on that rubric, instrument TSPT, build a defensible amortization story, and avoid the most expensive failure mode in AI enablement — “give everyone the tools” — by building the pond first.
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
We stopped equipping humans and started automating their worst hour. Instead of teaching people to fish, we built the pond: precise background agents pointed at repeated, skill-independent time-sinks where ROI was measurable on day one. The first build was a bug-triage agent wired to GitHub, Jira, and Notion; every agent that followed was measured the same way — Time Saved Per Task (TSPT) per week — against an amortized cost model (build + inference + maintenance vs. hours saved × loaded hourly cost).
Structurally, the change was moving from a portfolio of per-person tool subscriptions to a portfolio of scoped, measurable background agents with a pre-build attribution conversation baked in. The decision rubric — Sharpness × Cadence × Skill Tax → Pond / Boat / Net / Wade — tells you when to automate fully, when to augment, when to keep a human in the loop, and when the math doesn’t work yet.
Attendees will leave able to: score their own candidate workflows on that rubric, instrument TSPT, build a defensible amortization story, and avoid the most expensive failure mode in AI enablement — “give everyone the tools” — by building the pond first.
Vishvesh Pandey
Vishvesh Pandey
ABOUT THE SPEAKER:
Most practitioners instrumenting AI agents in production analytics workflows start by tracking token consumption — and discover too late that token spend tells you what you spent, not whether it was worth it. This talk walks through the operational reframe underway in financial services and adjacent regulated industries: from cost-per-token to cost-per-resolved-task as the primary economic unit, with call-level attribution as the foundation that makes everything else possible.
I’ll cover four cost engineering patterns that hold up in production analytics environments: (1) instrumentation at the call level — what to log, where it lives, who owns it; (2) model tiering and routing — when a frontier model earns its premium and when a smaller fine-tuned model wins; (3) the output-token asymmetry that makes structured outputs and chain-of-thought hygiene more economically consequential than they appear; and (4) cost containment as an engineering practice with runtime guardrails, not a finance review after the fact.
I’ll also be honest about what doesn’t generalize: caching strategies that worked for one workload and broke another, routing logic that added more latency than it saved, and the political dimension of asking analysts to justify their agent spend. Attendees will leave with a maturity model for agent cost engineering, a concrete instrumentation checklist they can apply next week, and a vocabulary for the cost conversation that bridges engineering and finance — because in regulated industries, that’s the conversation that decides whether agents scale or stall.
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
Most practitioners instrumenting AI agents in production analytics workflows start by tracking token consumption — and discover too late that token spend tells you what you spent, not whether it was worth it. This talk walks through the operational reframe underway in financial services and adjacent regulated industries: from cost-per-token to cost-per-resolved-task as the primary economic unit, with call-level attribution as the foundation that makes everything else possible.
I’ll cover four cost engineering patterns that hold up in production analytics environments: (1) instrumentation at the call level — what to log, where it lives, who owns it; (2) model tiering and routing — when a frontier model earns its premium and when a smaller fine-tuned model wins; (3) the output-token asymmetry that makes structured outputs and chain-of-thought hygiene more economically consequential than they appear; and (4) cost containment as an engineering practice with runtime guardrails, not a finance review after the fact.
I’ll also be honest about what doesn’t generalize: caching strategies that worked for one workload and broke another, routing logic that added more latency than it saved, and the political dimension of asking analysts to justify their agent spend. Attendees will leave with a maturity model for agent cost engineering, a concrete instrumentation checklist they can apply next week, and a vocabulary for the cost conversation that bridges engineering and finance — because in regulated industries, that’s the conversation that decides whether agents scale or stall.
AI Sovereignty
Anne Griffin
Anne Griffin
ABOUT THE SPEAKER:
When Hugging Face went to investigate this summer’s OpenAI containment breach, the commercial models they reached for refused to help, because their safety guardrails couldn’t tell defending from attacking apart. Instead, Hugging Face used an open weight model on their own infrastructure and finished the forensics in hours with none of the attacker data leaving their environment. It’s the perfect example of how open weight models let you control your guardrails, privacy, governance, latency, and reliability.
More companies are asking if they should use and self host open weight models, and this talk will walk through a framework to determine when it does and doesn’t make sense. By the end of this talk, attendees will understand what the benefits and drawbacks are of open weight models, and how they can bring both technical and business value.
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
When Hugging Face went to investigate this summer’s OpenAI containment breach, the commercial models they reached for refused to help, because their safety guardrails couldn’t tell defending from attacking apart. Instead, Hugging Face used an open weight model on their own infrastructure and finished the forensics in hours with none of the attacker data leaving their environment. It’s the perfect example of how open weight models let you control your guardrails, privacy, governance, latency, and reliability.
More companies are asking if they should use and self host open weight models, and this talk will walk through a framework to determine when it does and doesn’t make sense. By the end of this talk, attendees will understand what the benefits and drawbacks are of open weight models, and how they can bring both technical and business value.
Hardware and Chips
Kalpesh Sutaria
Kalpesh Sutaria
ABOUT THE SPEAKER:
A leaderboard win is not a product. Nemotron Retriever’s models rank #1 on RTEB — but topping a benchmark and running efficiently inside a customer’s production are two very different problems. This talk is the engineering story of closing that gap.
I’ll walk through our journey rebuilding the Retriever inference stack in Rust with a single obsession: treating state-of-the-art performance as a first-class product requirement, not a post-hoc optimization. We’ll get concrete about the decisions that mattered — what we measured, where we spent effort, the throughput, memory, and footprint wins we chased, and how dramatic efficiency gains unlocked deployment scenarios (including edge and on-device) that simply weren’t possible before. I’ll also share what didn’t work, the tension between shipping fast and building durable, and how we leaned on AI-assisted development to move faster than the roadmap assumed.
You’ll leave with a mental model for treating inference performance as a product and go-to-market lever — and a practical playbook for taking research-grade models to production without leaving speed, cost, or reach on the table.
Key takeaways:
- Why “performance is the product”: efficiency is an adoption and go-to-market lever, not a cost center.
- Concrete decisions from a Rust-based inference rebuild — what to measure, where to optimize, and how to know it worked.
- How latency and footprint reductions open entirely new deployment surfaces (edge, on-device) for the same models.
- Navigating the research-to-production seam: turning SoTA benchmark results into reliable, cheap, fast serving customers can bet on.
Who should attend / level? AI engineers, ML platform/infrastructure engineers, and engineering leaders who build, deploy, or operate inference and serving systems. Intermediate.
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
A leaderboard win is not a product. Nemotron Retriever’s models rank #1 on RTEB — but topping a benchmark and running efficiently inside a customer’s production are two very different problems. This talk is the engineering story of closing that gap.
I’ll walk through our journey rebuilding the Retriever inference stack in Rust with a single obsession: treating state-of-the-art performance as a first-class product requirement, not a post-hoc optimization. We’ll get concrete about the decisions that mattered — what we measured, where we spent effort, the throughput, memory, and footprint wins we chased, and how dramatic efficiency gains unlocked deployment scenarios (including edge and on-device) that simply weren’t possible before. I’ll also share what didn’t work, the tension between shipping fast and building durable, and how we leaned on AI-assisted development to move faster than the roadmap assumed.
You’ll leave with a mental model for treating inference performance as a product and go-to-market lever — and a practical playbook for taking research-grade models to production without leaving speed, cost, or reach on the table.
Key takeaways:
- Why “performance is the product”: efficiency is an adoption and go-to-market lever, not a cost center.
- Concrete decisions from a Rust-based inference rebuild — what to measure, where to optimize, and how to know it worked.
- How latency and footprint reductions open entirely new deployment surfaces (edge, on-device) for the same models.
- Navigating the research-to-production seam: turning SoTA benchmark results into reliable, cheap, fast serving customers can bet on.
Who should attend / level? AI engineers, ML platform/infrastructure engineers, and engineering leaders who build, deploy, or operate inference and serving systems. Intermediate.
Poonam Lamba
Poonam Lamba
ABOUT THE SPEAKER:
We redesigned distributed GPU orchestration for RL post-training and batch inference in the open-source llm-d platform. Structurally, we replaced static GPU/TPU locking with a three-tier co-operative time-slicing system:
- Application Layer: Workloads signal phase boundaries (rollouts, training, batch inference) via explicit acquire() and yield() APIs.
- Cluster Orchestrator: Manages lock queues to dynamically interleave complementary jobs onto shared hardware during idle phases.
- Node Snapshot Agent: Executes fast sub-second state swaps between GPU/TPU VRAM and host DRAM, enabling instant context switching without container restarts.
Attendees will walk away with:
- Drive 70%+ GPU/TPU Utilization: Understand how time-slicing reclaims idle hardware during RL loops and batch inference without impacting convergence.
- Architect Rapid Memory Swapping: Apply VRAM-to-DRAM snapshotting strategies for ultra-fast GPU/TPU context switching.
- Deploy on Kubernetes: Configure llm-d and K8s orchestrators to interleave RL and batch inference on shared clusters.
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
We redesigned distributed GPU orchestration for RL post-training and batch inference in the open-source llm-d platform. Structurally, we replaced static GPU/TPU locking with a three-tier co-operative time-slicing system:
- Application Layer: Workloads signal phase boundaries (rollouts, training, batch inference) via explicit acquire() and yield() APIs.
- Cluster Orchestrator: Manages lock queues to dynamically interleave complementary jobs onto shared hardware during idle phases.
- Node Snapshot Agent: Executes fast sub-second state swaps between GPU/TPU VRAM and host DRAM, enabling instant context switching without container restarts.
Attendees will walk away with:
- Drive 70%+ GPU/TPU Utilization: Understand how time-slicing reclaims idle hardware during RL loops and batch inference without impacting convergence.
- Architect Rapid Memory Swapping: Apply VRAM-to-DRAM snapshotting strategies for ultra-fast GPU/TPU context switching.
- Deploy on Kubernetes: Configure llm-d and K8s orchestrators to interleave RL and batch inference on shared clusters.
Fuzail Khan
Fuzail Khan
ABOUT THE SPEAKER:
The ranking and recommendation systems landscape is being transformed in the generative era. This talk reports on the experience of building and shipping the training infrastructure behind the first generative recommender in production at Meta, covering both stages of the recipe: pre-training to acquire the generative capability and post-training RL to align generation with the ranking objective.
Generative recommenders bring distinct challenges to end-to-end performance and scalability. This primarily arises from a mixed architecture that consists of both recommender-native large sparse embedding tables and LLM-based decoders. This then generates semantic IDs translating to real-world use cases in production such as finding the right advertisement for a given user. We inherit the communication profile of a sparse recommender as well as the autoregressive nature of a large language model for which an end-to-end systems blueprint simply does not exist.
For pre-training, we make the significant change from discriminative to generative recommendation. We talk about feature processing, data loading and the user-modeling path while building the LLM decoder and semantic ID tokenization on top while addressing real-world productionization challenges and end-to-end LLM performance analysis.
Reinforcement learning was implemented as an extensible framework where rewards, losses, RL algorithms, reference models and generation strategies are pluggable. We’ll walk through the reward design and the end-to-end post-training pipeline that prioritizes scalability at the production scale at Meta.
We’ll close with end-to-end optimization work – redundant computation elimination, specialized kernels for very short sequences, dense-sparse pipelining and hybrid embedding placement that satisfies large-scale performance requirements in compute and throughput.
We believe the key output from this session is leaving attendees with a crisp understanding of generative recommendation systems in large-scale production. We want to ensure we stay practical and address real-world production challenges that are relevant to developers and builders in AI infrastructure today. We do this by diving into the technicalities of enabling a large-scale generative retrieval system as well as concrete performance optimizations that enable the shift from a research prototype to a highly-optimized production system critical to revenue. We also feel the training systems and infrastructure-side of LLMs in production is rarely addressed at least relative to model architecture and quality, and we hope to fill that gap with this session.
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
The ranking and recommendation systems landscape is being transformed in the generative era. This talk reports on the experience of building and shipping the training infrastructure behind the first generative recommender in production at Meta, covering both stages of the recipe: pre-training to acquire the generative capability and post-training RL to align generation with the ranking objective.
Generative recommenders bring distinct challenges to end-to-end performance and scalability. This primarily arises from a mixed architecture that consists of both recommender-native large sparse embedding tables and LLM-based decoders. This then generates semantic IDs translating to real-world use cases in production such as finding the right advertisement for a given user. We inherit the communication profile of a sparse recommender as well as the autoregressive nature of a large language model for which an end-to-end systems blueprint simply does not exist.
For pre-training, we make the significant change from discriminative to generative recommendation. We talk about feature processing, data loading and the user-modeling path while building the LLM decoder and semantic ID tokenization on top while addressing real-world productionization challenges and end-to-end LLM performance analysis.
Reinforcement learning was implemented as an extensible framework where rewards, losses, RL algorithms, reference models and generation strategies are pluggable. We’ll walk through the reward design and the end-to-end post-training pipeline that prioritizes scalability at the production scale at Meta.
We’ll close with end-to-end optimization work – redundant computation elimination, specialized kernels for very short sequences, dense-sparse pipelining and hybrid embedding placement that satisfies large-scale performance requirements in compute and throughput.
We believe the key output from this session is leaving attendees with a crisp understanding of generative recommendation systems in large-scale production. We want to ensure we stay practical and address real-world production challenges that are relevant to developers and builders in AI infrastructure today. We do this by diving into the technicalities of enabling a large-scale generative retrieval system as well as concrete performance optimizations that enable the shift from a research prototype to a highly-optimized production system critical to revenue. We also feel the training systems and infrastructure-side of LLMs in production is rarely addressed at least relative to model architecture and quality, and we hope to fill that gap with this session.
Balaji Varadarajan
Balaji Varadarajan
ABOUT THE SPEAKER:
We will start with two diagnostic tools: Little’s Law which translates your QPS and target latency into the concurrency and batch size your system actually needs and the roofline model which tells you whether prefill (compute-bound) or decode (memory-bound) is your real constraint before you reach for a fix..
Will walk through the levers that move the needle most – MoE vs. dense architecture and what that does to your communication pattern FP8/INT4 quantization and where it costs you accuracy, TP/EP/DP parallelism and when each is the right tool and batching to the B_sat knee before compute-bound latency hockey-sticks..
Will cover concrete tuning playbooks for two workload types:
- latency-sensitive traffic like chat and agents where you’re optimizing TTFT and P99 ITL with chunked prefill and modest batch sizes
- Throughput-sensitive traffic like batch summarization and evals where you push past B_sat and let queuing work in your favor.
Most production systems live in both worlds at once, so we’ll need a balanced configuration to tackle both.
You will leave with a 7-step decision framework – characterize your workload, pick your model and quantization, benchmark on candidate hardware, find the knee, size your deployment with Little’s Law, calculate total cost of ownership, and plan for autoscaling that you can apply to your own inference stack..
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
We will start with two diagnostic tools: Little’s Law which translates your QPS and target latency into the concurrency and batch size your system actually needs and the roofline model which tells you whether prefill (compute-bound) or decode (memory-bound) is your real constraint before you reach for a fix..
Will walk through the levers that move the needle most – MoE vs. dense architecture and what that does to your communication pattern FP8/INT4 quantization and where it costs you accuracy, TP/EP/DP parallelism and when each is the right tool and batching to the B_sat knee before compute-bound latency hockey-sticks..
Will cover concrete tuning playbooks for two workload types:
- latency-sensitive traffic like chat and agents where you’re optimizing TTFT and P99 ITL with chunked prefill and modest batch sizes
- Throughput-sensitive traffic like batch summarization and evals where you push past B_sat and let queuing work in your favor.
Most production systems live in both worlds at once, so we’ll need a balanced configuration to tackle both.
You will leave with a 7-step decision framework – characterize your workload, pick your model and quantization, benchmark on candidate hardware, find the knee, size your deployment with Little’s Law, calculate total cost of ownership, and plan for autoscaling that you can apply to your own inference stack..
Recursive Self-Improvement (RSI)
Vashishtha Patil
Vashishtha Patil
ABOUT THE SPEAKER:
Autonomous research agents that propose, implement, and refine ML solutions are priced out by the loop, not the model. Published agents assume a frontier model drives every step, so cost scales with trajectory length on runs that are long by design. Teams either cap the horizon — removing the thing that made it work — or don’t run it at all.
Hypothesis: a small open-weight model runs the loop end to end, with a frontier advisor called in at those points on a metered budget, holding most of the baseline’s result quality at a fraction of its cost. The open question is whether the escalation trigger is reliable enough to justify the calls it buys — our first version often bought advice the loop would have reached on its own.
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
Autonomous research agents that propose, implement, and refine ML solutions are priced out by the loop, not the model. Published agents assume a frontier model drives every step, so cost scales with trajectory length on runs that are long by design. Teams either cap the horizon — removing the thing that made it work — or don’t run it at all.
Hypothesis: a small open-weight model runs the loop end to end, with a frontier advisor called in at those points on a metered budget, holding most of the baseline’s result quality at a fraction of its cost. The open question is whether the escalation trigger is reliable enough to justify the calls it buys — our first version often bought advice the loop would have reached on its own.
Agent Safety and Security
Siddharth Jain
Siddharth Jain
ABOUT THE SPEAKER:
Teams often give an agent a service credential, add a human approval step, and call the workflow governed. That design breaks down when the agent can revise payloads, retry writes, chain tools, or act across systems with different permission models. The result is an accountability gap: the organization can see that a service account acted, but not necessarily who authorized the business intent, which payload was approved, whether a retry duplicated work, or what changed between proposal and execution.
This session presents a production control model for agent workflows that make consequential writes. It shows how to separate the business-intent identity from the concrete operation; classify tools by impact; issue short-lived, least-privilege credentials only after validation; bind human approval to a canonical payload hash and policy version; execute through controlled services with idempotency keys; and reconcile external state before declaring success. It also covers the evidence record needed to answer four operational questions: who requested the action, what the agent proposed, what policy and human approved, and what the downstream system actually did.
The talk focuses on failure modes that appear after the demo works: stale approvals, overbroad agent permissions, duplicate writes after retries, silent policy changes, and actions whose outcome is unknown. Attendees will leave with a lifecycle they can map to their own platform, concrete interfaces between model output and deterministic controls, and an audit schema that supports incident response, compliance review, and day-to-day operations without turning every agent into a bespoke security project.
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
Teams often give an agent a service credential, add a human approval step, and call the workflow governed. That design breaks down when the agent can revise payloads, retry writes, chain tools, or act across systems with different permission models. The result is an accountability gap: the organization can see that a service account acted, but not necessarily who authorized the business intent, which payload was approved, whether a retry duplicated work, or what changed between proposal and execution.
This session presents a production control model for agent workflows that make consequential writes. It shows how to separate the business-intent identity from the concrete operation; classify tools by impact; issue short-lived, least-privilege credentials only after validation; bind human approval to a canonical payload hash and policy version; execute through controlled services with idempotency keys; and reconcile external state before declaring success. It also covers the evidence record needed to answer four operational questions: who requested the action, what the agent proposed, what policy and human approved, and what the downstream system actually did.
The talk focuses on failure modes that appear after the demo works: stale approvals, overbroad agent permissions, duplicate writes after retries, silent policy changes, and actions whose outcome is unknown. Attendees will leave with a lifecycle they can map to their own platform, concrete interfaces between model output and deterministic controls, and an audit schema that supports incident response, compliance review, and day-to-day operations without turning every agent into a bespoke security project.
WHAT YOU’LL LEARN:
- Do not let the model hold broad standing credentials; mint scoped identity for each authorized operation.
- Bind approval to material fields, a canonical payload hash, and the policy version so edits invalidate stale consent.
- Treat every consequential write as a durable state machine with explicit recovery paths, not as a single tool call.
- Separate retry from reconciliation so an ambiguous timeout does not become a duplicate action.
- Design the audit record as an operating interface: it should reconstruct who requested, what was proposed, what was approved, what executed, and what the downstream system recorded.
PREREQUISITE KNOWLEDGE
Agent Memory Architectures
AI-Assisted Software Engineering
Technical / Engineering
Zachary Hamilton
Zachary Hamilton
ABOUT THE SPEAKER:
Building agentic software requires more than adding an LLM to the traditional software development lifecycle. When behavior becomes non-deterministic, teams need a new feedback loop connecting what happens in production to how they evaluate, debug, and improve their systems. This talk will show how to build that loop: identifying real failure modes from production traces and human feedback, turning them into reproducible datasets and evaluations, defining success using both technical and business outcomes, and using automated and human judges to measure whether changes actually improve the system. We’ll also explore where automated evaluation breaks down, how to calibrate LLM-as-a-judge, and how continuous evaluation helps teams detect regressions and emerging behavior. Attendees will leave with a practical framework for moving from production > failure discovery > datasets > evals > iteration > production, turning the development of agentic systems into a measurable engineering discipline.
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
Building agentic software requires more than adding an LLM to the traditional software development lifecycle. When behavior becomes non-deterministic, teams need a new feedback loop connecting what happens in production to how they evaluate, debug, and improve their systems. This talk will show how to build that loop: identifying real failure modes from production traces and human feedback, turning them into reproducible datasets and evaluations, defining success using both technical and business outcomes, and using automated and human judges to measure whether changes actually improve the system. We’ll also explore where automated evaluation breaks down, how to calibrate LLM-as-a-judge, and how continuous evaluation helps teams detect regressions and emerging behavior. Attendees will leave with a practical framework for moving from production > failure discovery > datasets > evals > iteration > production, turning the development of agentic systems into a measurable engineering discipline.
Workshops
Upal Saha
Upal Saha
ABOUT THE SPEAKER:
This is a build-and-break lab for engineers who already know that a demo is not a system. In the first fifteen minutes every attendee stands up a working extraction pipeline against a real invoice, with no schema authoring, and gets structured JSON back. Then we spend an hour breaking it the way production does, and fixing each break with a pattern that transfers to any stack.
Break one: the wrong document. We feed a bill of lading glued to an invoice into the invoice pipeline and watch it confidently produce garbage. Fix: classify before you extract, and make the graph deterministic even though every step inside it is a model. Break two: the answer that is probably right. Models do not tell you how confident they are, so we compute per-field confidence, find the fields that fall below 95%, and route those, and only those, to a human whose correction feeds back into the system. Break three: the messy string. “10 cases organic gala apples, 88 ct” has to become one SKU; we show why canonicalization is its own step, how to score matches, and where to set the threshold for review. Break four: the hostile input. Everyone runs an image whose pixels read “ignore all previous instructions” and we discuss, with the result on screen, what it means to treat inbound data as data rather than control. We close by labeling a handful of outputs and running a regression test between two versions of the pipeline, because evaluation is a loop, not a phase.
Attendees leave with a running pipeline in their own account and five patterns they can apply on Monday regardless of vendor: classify-then-extract, confidence scoring for models that lack it, threshold-based exception routing with feedback, semantic security checks on inbound data, and versioned regression testing. The platform used in the room is ours; the lessons are not.
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
This is a build-and-break lab for engineers who already know that a demo is not a system. In the first fifteen minutes every attendee stands up a working extraction pipeline against a real invoice, with no schema authoring, and gets structured JSON back. Then we spend an hour breaking it the way production does, and fixing each break with a pattern that transfers to any stack.
Break one: the wrong document. We feed a bill of lading glued to an invoice into the invoice pipeline and watch it confidently produce garbage. Fix: classify before you extract, and make the graph deterministic even though every step inside it is a model. Break two: the answer that is probably right. Models do not tell you how confident they are, so we compute per-field confidence, find the fields that fall below 95%, and route those, and only those, to a human whose correction feeds back into the system. Break three: the messy string. “10 cases organic gala apples, 88 ct” has to become one SKU; we show why canonicalization is its own step, how to score matches, and where to set the threshold for review. Break four: the hostile input. Everyone runs an image whose pixels read “ignore all previous instructions” and we discuss, with the result on screen, what it means to treat inbound data as data rather than control. We close by labeling a handful of outputs and running a regression test between two versions of the pipeline, because evaluation is a loop, not a phase.
Attendees leave with a running pipeline in their own account and five patterns they can apply on Monday regardless of vendor: classify-then-extract, confidence scoring for models that lack it, threshold-based exception routing with feedback, semantic security checks on inbound data, and versioned regression testing. The platform used in the room is ours; the lessons are not.
Rajiv Shah
Rajiv Shah
ABOUT THE SPEAKER:
You can start a simple agent with a model and a prompt. But you soon realize that improving it requires adjusting what the model can see, what it can do, how it remembers progress, which model handles each step, and what evidence allows the work to stop. Harness engineering is how those pieces become one working system.
This workshop uses coding agents to make that system concrete. Participants will work through six decisions: choosing a harness, designing tools and retrieval, placing context and memory, routing models and reasoning, controlling goals and validation, and deciding when a task benefits from multiple agents. Through hands-on experiments, we will change these settings and inspect what happens in the trace and final result.
Each exercise is paired with an overview of current research on agent harnesses. Together, the research and experiments show how the major components interact and where each approach reaches its limits. Participants will leave with a practical way to reason about the whole agent, not just its model.
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
You can start a simple agent with a model and a prompt. But you soon realize that improving it requires adjusting what the model can see, what it can do, how it remembers progress, which model handles each step, and what evidence allows the work to stop. Harness engineering is how those pieces become one working system.
This workshop uses coding agents to make that system concrete. Participants will work through six decisions: choosing a harness, designing tools and retrieval, placing context and memory, routing models and reasoning, controlling goals and validation, and deciding when a task benefits from multiple agents. Through hands-on experiments, we will change these settings and inspect what happens in the trace and final result.
Each exercise is paired with an overview of current research on agent harnesses. Together, the research and experiments show how the major components interact and where each approach reaches its limits. Participants will leave with a practical way to reason about the whole agent, not just its model.
Chris Alexiuk
Chris Alexiuk
ABOUT THE SPEAKER:
In this workshop we’ll stand up NemoClaw end to end: install the reference stack, get OpenClaw running inside the OpenShell sandbox, configure inference routing, and lock down a network policy that survives a multi-hour agent session. We’ll walk through the blueprint, the CLI, and the approval flow, then run a real long-lived agent against it and break things on purpose so you know what the layers actually catch.
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
In this workshop we’ll stand up NemoClaw end to end: install the reference stack, get OpenClaw running inside the OpenShell sandbox, configure inference routing, and lock down a network policy that survives a multi-hour agent session. We’ll walk through the blueprint, the CLI, and the approval flow, then run a real long-lived agent against it and break things on purpose so you know what the layers actually catch.
Virtual Day
Jim Allen Wallace
Jim Allen Wallace
ABOUT THE SPEAKER:
Instacart’s ad-serving feature store served features to real-time inference from a multi-hundred-node managed Valkey deployment split across multiple clusters. The team had built a proxy layer, a querying SDK, and a compact storage format on top of it, and still hit three limits: instability during any cluster mutation, tail latency from read fan-out, and cost that got worse when they split clusters to manage the first two.
The structural change was to stop scaling out and scale up. A multi-threaded engine let the team run much larger instances, so hundreds of nodes became roughly 100, each inference request touched far fewer shards, and average and P99 latency fell 50%. Because the proxy and SDK already hid the datastore from ML engineers, the swap required no application changes and tens of terabytes moved without anyone above the storage layer noticing.
We were the engine vendor, and the first pass did not reach the latency the team expected. Getting there took client-side tuning on their end and engine changes from our engineering team, including to the compactor. Validation ran side by side: identical backfill, traffic shifted from 0% to 100%, gates at each percentile.
Attendees will leave able to test a datastore against the fan-out shape of their own feature store rather than an ops/sec number, trace how engine concurrency sets instance size, which sets fan-out, which sets P99, and run a datastore migration under live traffic with per-percentile gates and a rollback path. They will also hear where the trade-off comes back as the cluster grows.
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
Instacart’s ad-serving feature store served features to real-time inference from a multi-hundred-node managed Valkey deployment split across multiple clusters. The team had built a proxy layer, a querying SDK, and a compact storage format on top of it, and still hit three limits: instability during any cluster mutation, tail latency from read fan-out, and cost that got worse when they split clusters to manage the first two.
The structural change was to stop scaling out and scale up. A multi-threaded engine let the team run much larger instances, so hundreds of nodes became roughly 100, each inference request touched far fewer shards, and average and P99 latency fell 50%. Because the proxy and SDK already hid the datastore from ML engineers, the swap required no application changes and tens of terabytes moved without anyone above the storage layer noticing.
We were the engine vendor, and the first pass did not reach the latency the team expected. Getting there took client-side tuning on their end and engine changes from our engineering team, including to the compactor. Validation ran side by side: identical backfill, traffic shifted from 0% to 100%, gates at each percentile.
Attendees will leave able to test a datastore against the fan-out shape of their own feature store rather than an ops/sec number, trace how engine concurrency sets instance size, which sets fan-out, which sets P99, and run a datastore migration under live traffic with per-percentile gates and a rollback path. They will also hear where the trade-off comes back as the cluster grows.
Jessica Garson Beauchemin
Jessica Garson Beauchemin
ABOUT THE SPEAKER:
Deploying AI applications often involves containers, GPU configuration, dependency management, and infrastructure scaling. Runpod Flash offers a simpler, code-first approach. It lets you define remote functions and hardware requirements in Python and run them on serverless GPUs.
In this talk, we’ll explore how Flash moves Python workloads from local development to cloud deployment, examine its underlying programming model, and build a GPU-backed endpoint. Along the way, we’ll discuss where Flash fits in the AI development stack, the problems it solves, and the trade-offs developers should consider. Attendees will leave with a practical understanding of how to turn local AI code into a scalable service without needing to become infrastructure experts.
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
Deploying AI applications often involves containers, GPU configuration, dependency management, and infrastructure scaling. Runpod Flash offers a simpler, code-first approach. It lets you define remote functions and hardware requirements in Python and run them on serverless GPUs.
In this talk, we’ll explore how Flash moves Python workloads from local development to cloud deployment, examine its underlying programming model, and build a GPU-backed endpoint. Along the way, we’ll discuss where Flash fits in the AI development stack, the problems it solves, and the trade-offs developers should consider. Attendees will leave with a practical understanding of how to turn local AI code into a scalable service without needing to become infrastructure experts.
Andy McMahon
Andy McMahon
ABOUT THE SPEAKER:
As organisations move from single models to connected agent systems, the challenge shifts to how agents interact, how they are evaluated, and how their behaviour can be monitored in real production environments.
This presentation explores what it takes to operationalise agentic AI in practice – from observability and evaluation to deploying agents safely and reliably across enterprise workflows.
Through real-world examples, we look at how teams are building responsible AgentOps frameworks to scale multi-agent systems in production while maintaining control, trust and measurable business impact.
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
As organisations move from single models to connected agent systems, the challenge shifts to how agents interact, how they are evaluated, and how their behaviour can be monitored in real production environments.
This presentation explores what it takes to operationalise agentic AI in practice – from observability and evaluation to deploying agents safely and reliably across enterprise workflows.
Through real-world examples, we look at how teams are building responsible AgentOps frameworks to scale multi-agent systems in production while maintaining control, trust and measurable business impact.
Yegor Denisov-Blanch
Yegor Denisov-Blanch
ABOUT THE SPEAKER:
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
Vicente Rubén Del Pino Ruiz
Vicente Rubén Del Pino Ruiz
ABOUT THE SPEAKER:
Single-shot evaluation, the shape every current AI eval framework ships, is structurally blind to the failures that take agents down in production: patience that runs out at turn six, users who abandon silently, tool calls that lose conversation state, partial successes that masquerade as wins. This talk argues for a different shape, drawn from the digital twin literature and how the approach is already applied in aerospace, autonomous vehicles, and civil engineering: simulate the population of users your agent will meet, run them against the agent, watch what breaks before any human sees it.
Three take aways:
- Why current AI agent evaluation cannot detect the failures that hit production. The structural reason prompt-and-grade testing (the shape every current framework ships) is blind to patience, abandonment, conversation state, and partial success.
- What a realistic synthetic user population looks like, drawn from the digital twin literature in other engineering industries. Four properties: continuous trait clusters instead of fixed persona archetypes; probabilistic dropout (real users walk away silently rather than say goodbye); three-way goal scoring (each conversation is scored “achieved,” “partially achieved,” or “not achieved,” instead of just pass/fail); and variable conversation length (each conversation ends when the persona’s patience runs out, not at a fixed turn count).
- How to wire population-scale evaluation into a release gate. Why run-over-run delta on a held-constant population is the regression-detection primitive AI agents actually need, and how to treat it as a governance lever rather than a research experiment.”
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
Single-shot evaluation, the shape every current AI eval framework ships, is structurally blind to the failures that take agents down in production: patience that runs out at turn six, users who abandon silently, tool calls that lose conversation state, partial successes that masquerade as wins. This talk argues for a different shape, drawn from the digital twin literature and how the approach is already applied in aerospace, autonomous vehicles, and civil engineering: simulate the population of users your agent will meet, run them against the agent, watch what breaks before any human sees it.
Three take aways:
- Why current AI agent evaluation cannot detect the failures that hit production. The structural reason prompt-and-grade testing (the shape every current framework ships) is blind to patience, abandonment, conversation state, and partial success.
- What a realistic synthetic user population looks like, drawn from the digital twin literature in other engineering industries. Four properties: continuous trait clusters instead of fixed persona archetypes; probabilistic dropout (real users walk away silently rather than say goodbye); three-way goal scoring (each conversation is scored “achieved,” “partially achieved,” or “not achieved,” instead of just pass/fail); and variable conversation length (each conversation ends when the persona’s patience runs out, not at a fixed turn count).
- How to wire population-scale evaluation into a release gate. Why run-over-run delta on a held-constant population is the regression-detection primitive AI agents actually need, and how to treat it as a governance lever rather than a research experiment.”
Matt Mazzarell
Matt Mazzarell
ABOUT THE SPEAKER:
“One of the most difficult problems every company faces is understanding its customers completely. Customer lifetime value, attrition risk, and purchase propensity are all solvable with AI/ML — but how do we combine these modeling scores to initiate the right action with the right customer at any point in time?
Agentic applications help us make the best possible decisions when interpreting complex, high-volume signals from our customers. An agentic application gives end users visuals that explain key insights, with an agent in the loop to ensure nothing is missed. Context is everything: when done correctly, the agent always has the appropriate understanding to build an action plan that improves customer health and profitability.
In this session, we’ll show you how to build agentic apps from ideation to a finished product that interacts with customers. You’ll take away practical tips for using agentic coding frameworks, curating complete customer data products, and building customer-facing agents — capped off with a live demo of Teradata’s Customer Lifetime Value Agentic App.”
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
“One of the most difficult problems every company faces is understanding its customers completely. Customer lifetime value, attrition risk, and purchase propensity are all solvable with AI/ML — but how do we combine these modeling scores to initiate the right action with the right customer at any point in time?
Agentic applications help us make the best possible decisions when interpreting complex, high-volume signals from our customers. An agentic application gives end users visuals that explain key insights, with an agent in the loop to ensure nothing is missed. Context is everything: when done correctly, the agent always has the appropriate understanding to build an action plan that improves customer health and profitability.
In this session, we’ll show you how to build agentic apps from ideation to a finished product that interacts with customers. You’ll take away practical tips for using agentic coding frameworks, curating complete customer data products, and building customer-facing agents — capped off with a live demo of Teradata’s Customer Lifetime Value Agentic App.”
Antonio Bustamante
Antonio Bustamante
ABOUT THE SPEAKER:
Everyone building on frontier models hits the same wall: 80% of the way there in a weekend, then an exponentially expensive climb toward the 99%+ that operational systems need. This talk is the honest map of that climb, from a team that now runs AI over millions of documents, images and videos a month for customers in logistics, fleet management, automotive and financial services who need the answer to be right every time.
We will walk through the ladder nobody budgets for: retries, then queues when the model is down for three hours, then rate limits, then discovering that 95% is not enough for transactional data, then discovering that the model cannot tell you how confident it is. We will show two failures from our own production history, a customer whose single engineer racked up $30K of usage in a month because nothing was watching, and our first churn, on 500-page reports with a hundred rows per page, a problem we still consider unsolved. And we will show what we changed: a harness that treats AI as a deterministic step inside durable workflows rather than as an open-ended agent, decisions expressed as a verified tree the model must traverse, algorithmic confidence scoring on top of models that provide none, routing anything under 95% to a human whose verdict feeds back into the system, semantic checks against instructions smuggled into the data, and, counterintuitively, encouraging customers to build their own independent monitoring of us. One customer’s users went from eight to nine hours a week on a task to about thirty minutes, and that number was measured by them, not by us.
Attendees will leave with five patterns they can apply Monday: confidence scoring for models that lack it, decision trees over open-ended prompts, exception routing with feedback loops, semantic security checks on inbound data, and customer-owned evaluation. They will also leave with a thesis we did not start with: chat is single-player AI; the next decade of software is ambient AI that runs the same process a hundred thousand times a day, unattended, and behaves the same way every time.
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
Everyone building on frontier models hits the same wall: 80% of the way there in a weekend, then an exponentially expensive climb toward the 99%+ that operational systems need. This talk is the honest map of that climb, from a team that now runs AI over millions of documents, images and videos a month for customers in logistics, fleet management, automotive and financial services who need the answer to be right every time.
We will walk through the ladder nobody budgets for: retries, then queues when the model is down for three hours, then rate limits, then discovering that 95% is not enough for transactional data, then discovering that the model cannot tell you how confident it is. We will show two failures from our own production history, a customer whose single engineer racked up $30K of usage in a month because nothing was watching, and our first churn, on 500-page reports with a hundred rows per page, a problem we still consider unsolved. And we will show what we changed: a harness that treats AI as a deterministic step inside durable workflows rather than as an open-ended agent, decisions expressed as a verified tree the model must traverse, algorithmic confidence scoring on top of models that provide none, routing anything under 95% to a human whose verdict feeds back into the system, semantic checks against instructions smuggled into the data, and, counterintuitively, encouraging customers to build their own independent monitoring of us. One customer’s users went from eight to nine hours a week on a task to about thirty minutes, and that number was measured by them, not by us.
Attendees will leave with five patterns they can apply Monday: confidence scoring for models that lack it, decision trees over open-ended prompts, exception routing with feedback loops, semantic security checks on inbound data, and customer-owned evaluation. They will also leave with a thesis we did not start with: chat is single-player AI; the next decade of software is ambient AI that runs the same process a hundred thousand times a day, unattended, and behaves the same way every time.
Christopher G. Potts
Christopher G. Potts
ABOUT THE SPEAKER:
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
Jazmia Henry
Jazmia Henry
ABOUT THE SPEAKER:
There’s a gap between researcher-crafted evaluation frameworks that capture model performance and benchmarks versus how end users actually use AI products. This gap is being exploited in ways that render traditional reward models useless.
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
There’s a gap between researcher-crafted evaluation frameworks that capture model performance and benchmarks versus how end users actually use AI products. This gap is being exploited in ways that render traditional reward models useless.
WHAT YOU’LL LEARN:
First, if your domain has computable ground truth, you do not need human annotators to build a reward signal. Second, scalar reward is a compression that loses information. Decomposing reward into multiple verifiable dimensions exposes failure modes that single-number metrics hide. Third, the exploitation gap is a distribution problem, not a labeling problem. Agents exploit the distance between training distribution and deployment reality, and deterministic verifiers narrow that gap directly.
PREREQUISITE KNOWLEDGE
Sandeep Bharadwaj Mannapur
Sandeep Bharadwaj Mannapur
ABOUT THE SPEAKER:
We caught our LLM quality problem the wrong way: through customer complaints. Three weeks of degraded responses had already shipped. Every monitoring signal we had was green the entire time because we were measuring the wrong things. Latency, error rate, embedding similarity: none of them captured what was actually happening, which was that response quality had quietly gotten worse in ways users noticed but our dashboards could not.
After that incident we rebuilt how we think about LLM observability. The core insight was that output quality is not the same as system health, and you cannot infer one from the other. We needed a separate signal layer built from production behavior: how users responded to answers, where downstream tasks broke down, where retrieval and response stopped agreeing with each other.
Getting this right took a few tries. A naive implementation fires constantly on normal LLM output variance. The real work was designing alerts that distinguish actual degradation from noise, and calibrating thresholds against real incident history rather than theoretical bounds.
You will leave with a taxonomy of LLM drift types and how each one shows up differently in production, the behavioral signals that actually correlate with quality degradation, and an alert pattern that catches real drift early without burying your team in false positives.
TALK TITLE:
TRACK:
Technical Level:
ABSTRACT:
We caught our LLM quality problem the wrong way: through customer complaints. Three weeks of degraded responses had already shipped. Every monitoring signal we had was green the entire time because we were measuring the wrong things. Latency, error rate, embedding similarity: none of them captured what was actually happening, which was that response quality had quietly gotten worse in ways users noticed but our dashboards could not.
After that incident we rebuilt how we think about LLM observability. The core insight was that output quality is not the same as system health, and you cannot infer one from the other. We needed a separate signal layer built from production behavior: how users responded to answers, where downstream tasks broke down, where retrieval and response stopped agreeing with each other.
Getting this right took a few tries. A naive implementation fires constantly on normal LLM output variance. The real work was designing alerts that distinguish actual degradation from noise, and calibrating thresholds against real incident history rather than theoretical bounds.
You will leave with a taxonomy of LLM drift types and how each one shows up differently in production, the behavioral signals that actually correlate with quality degradation, and an alert pattern that catches real drift early without burying your team in false positives.
Agenda
We will release it soon.
This agenda is still subject to changes.
Join free virtual sessions October 6–7, then meet us in Austin for in-person case studies, workshops, and expo October 8–9