
Building Production-Ready AI Systems: Security, AgentOps & LLM Evaluation
About this event
Building Production-Ready AI Systems: Security, AgentOps & LLM Evaluation
๐
July 23, 2026 | 5:30 PM โ 6:30 PM PT
๐ฅ Virtual Event
๐ Register on Microsoft Reactor:
https://aka.ms/ProdReady723/m
AI is no longer just about building models.
Today's AI applications require security, evaluation, observability, governance, and continuous improvement to succeed in production.
Join experts from Microsoft and Amazon as they share practical lessons from building, securing, evaluating, and operating AI-powered systems at scale.
As AI applications evolve from prototypes into real-world products, engineering teams face new challenges:
โข How do you protect AI systems from prompt injection, tool abuse, and emerging agent threats?
โข How do you evaluate whether an LLM is actually performing well in production?
โข How do you monitor, debug, and improve AI agents over time?
โข How do you fine-tune models for domain-specific workflows and measurable business impact?
This session brings together three critical pillars of modern AI engineering:
๐ AI Security
โ๏ธ AgentOps & Observability
๐ LLM Evaluation & Fine-Tuning
Featured Talks
Securing the AI Stack โ From Models to Agents to Infrastructure
Kriti Faujdar
Senior Product Manager, Microsoft Security AI Research
Learn a practical defense-in-depth framework for AI systems, covering prompt injection, jailbreaks, tool and MCP security, memory poisoning, sandboxing, secret management, and infrastructure-level protections.
AgentOps in the Open: Tools for Building, Testing, and Trusting AI Agents
Debjyoti Paul
Applied Scientist, Amazon
Explore the emerging AgentOps ecosystem and learn how teams are tracing agent behavior, evaluating tool calls, monitoring failures, testing prompts and workflows, and building feedback loops for continuous improvement.
Topics include Langfuse, OpenTelemetry, DeepEval, RAGAS, prompt versioning, testing frameworks, and production observability.
LLM-Driven Merge Conflict Resolution
Advitya Gemawat
Machine Learning Engineer, Microsoft
Discover how custom LLM evaluations and Azure OpenAI fine-tuning were used to build an AI-powered merge conflict resolver for one of the world's largest software codebases. Learn practical lessons from deploying LLM-powered developer tools, designing evaluation frameworks, and adapting models to domain-specific workflows.
What You'll Learn
โ
Security best practices for AI agents and applications
โ
How to evaluate LLMs beyond traditional benchmarks
โ
AgentOps tools and observability techniques
โ
Azure OpenAI fine-tuning workflows
โ
Real-world lessons from Microsoft and Amazon
โ
Practical approaches for building trustworthy AI systems
Who Should Attend?
โข Software Engineers
โข Machine Learning Engineers
โข AI Engineers
โข Data Scientists
โข Platform Engineers
โข Product Managers
โข Anyone building AI agents, copilots, RAG systems, or LLM-powered applications
Whether you're experimenting with AI agents or deploying production AI systems, you'll leave with practical frameworks, tools, and engineering insights you can apply immediately.
Hosted by PyData Seattle ร Microsoft Reactor
Source: meetup