Brainberg
AI Inference in the Real World: From Kubernetes to the Far Edge
AI Integration & ApplicationMeetupFree

AI Inference in the Real World: From Kubernetes to the Far Edge

Tue 27 Oct · 17:00
Amsterdam, 🇳🇱 Netherlands
50–200 attendees
AWS Amsterdam · Mr.Treublaan 7, 1097 DP Amsterdam, Netherlands

About this event

AI workloads don’t all belong on the same model, hardware, or infrastructure.

For our 13th AI Native Netherlands meetup, Christian Melendez from AWS will look at when smaller models and CPUs make more sense than defaulting to large GPU-backed LLMs.

William Rizzo from Mirantis will take that question to the far edge: running LLM inference on immutable, disconnected clusters where connectivity is unreliable and remote access can’t be assumed.

Together, the talks explore one practical question:

What should run where, on what hardware, and how do you keep it reliable in production?

A huge thank you to our friends at AWS for hosting us at their Amsterdam office. Food and drinks will be provided!

We'll cover:

  • When AI workloads actually need GPUs — and when CPUs and smaller models are the better fit.
  • How to combine Small Language Models and larger LLMs without sending every task to the most expensive model.
  • What changes when inference moves from the datacenter to factories, vehicles, retail sites, and other edge environments.
  • How immutable, image-based infrastructure can support updates, rollbacks, and reliability across disconnected fleets.
  • The real-world trade-offs, failure modes, and production lessons behind both approaches.

Speaker 1: Christian Melendez (AWS)
Christian Melendez is a Principal Specialist Solutions Architect at AWS. He helps the region's largest enterprises build efficient, resilient AI and cloud-native workloads on Kubernetes, with a focus on compute efficiency, cost optimisation, and autoscaling at scale. Author of Kubernetes Autoscaling and creator of Karpenter Blueprints, a best-practices repository that grew Karpenter adoption 16.5x across EMEA, he also built Slemify, an open-source framework for fine-tuning and serving Small Language Models on Kubernetes. A regular speaker at AWS re:Invent, KCDs / CNDs, ContainerDays, and others, Christian focuses on making infrastructure simple, observable, and cost-effective at scale.

Talk: The Right Tool for the Right Token: When to Reach for a CPU or a GPU
While hard reasoning problems require an LLM and a GPU, other tasks can be solved with simpler means. Classifying, routing, embedding, and reranking are structured, high-frequency jobs that a fine-tuned small model (SLM) can handle on CPU at predictable cost. In this session, we will demo a pipeline where each task runs on the best fitting infrastructure.

This will include a real CPU-first agent on Amazon EKS running an SLM, and calling an LLM when needed. We will share our learnings from building this pipeline, including how we measured price/performance, its limitations, and which of our design assumptions we found to be wrong.

You will leave with confidence to choose the right tool, and an open-source reference architecture to validate your assumptions.

Speaker 2: William Rizzo (Mirantis)
William Rizzo is Global Field CTO at Mirantis, where he helps organisations design, build, and run platform engineering, edge, and AI infrastructure initiatives. His career spans engineering, pre-sales, product ownership, and consulting across high-performance computing, storage, and distributed systems. A CNCF and Linkerd Ambassador and a Kairos maintainer, he's a regular speaker at KubeCon on platform engineering and building resilient internal developer platforms.

Talk: Inference at the Far Edge: Running LLMs on Immutable, Disconnected Clusters
Most inference architectures assume reliable connectivity, abundant infrastructure, and an operator who can access the system when something goes wrong.

At the edge, those assumptions disappear.

William will show how LLM inference can run on immutable, read-only edge nodes across factories, retail sites, vehicles, and remote facilities. He’ll cover how the model, runtime, and GPU drivers can be packaged into atomically updated images, how A/B upgrades and rollbacks work without a remote shell, and how inference can stay healthy across disconnected fleets.

He’ll also share the reference architecture, production failure modes, and the areas where edge inference still remains difficult.

The architecture is built entirely on CNCF and open-source projects.

Agenda:
18:00 — Arrival, food & drinks
18:45 — Talk #1 | Christian Melendez (AWS)
19:30 — Talk #2 | William Rizzo (Mirantis)
20:15 — Open conversation, networking & more drinks
21:00 — Wrapping up

What to bring:
Just curiosity and questions. If you're working on applied AI, MLOps, model serving, or the cost and infrastructure behind it, we'd love to hear how you're approaching it.

Who this is for:
AI/ML engineers, platform engineers, MLOps specialists, SREs, infrastructure engineers, architects, engineering leaders, and anyone working on production AI infrastructure.

Where to find us:
AWS Office: Mr. Treublaan 7, 1097 JS Amsterdam

Source: meetup