Agenda

Plan your day

Select a session to read the abstract and meet the speakers.

All times are local to Budapest.

Registration

–

Welcome + Opening Remarks

–

– Keynote Room 1

  • László Kollár

    Technology Manager, SMP Solutions

– Keynote Room 1

  • Piotr Trębacz

    Senior DevOps Manager, IGT

Break

–

– Room 1

  • Zsolt Varga

    CTO, Riptides

  • Nándor Krácser

    CISO, Riptides

– Room 2

  • Richard Kovacs

    CTO, HariKube

– Room 1

  • Viktor Nagy

    Fractional Head of Product and R&D, Clearvis PMS Kft

– Room 2

  • Gergely Daroczi

    Project Lead, Spare Cores

  • Attila Nagy

    cloud architect, Spare Cores Kft

– Room 1

  • Márton Enyedi

    Devops Engineer, Origoss Solutions

– Room 2

  • Dominik Táskai

    Senior Software Engineer, Genesys

Lunch

– · Room 1

Lunch

– · Room 2

– Room 1

  • Adam Hegedus

    Engineering Manager, Platform, Zocks

– Room 2

  • Hajnal Máté

    Senior Machine Learning Operations Engineer, Aumovio SE

– Room 1

  • Paul Power

    Software Engineer, Red Hat

– Room 2

  • Dmytro Kozlov

    Software engineer, VictoriaMetrics

– Room 1

  • Janos Binder

    Staff Engineer, SAP SE

– Room 2

  • Gergő Huszty

    Cloud infrastructure engineer, architect, IBM

Break

–

– Room 1

  • Bence Csati

    Senior Software Engineer, Axoflow

– Room 2

  • Xavier Avrillier

    Solutions architect, Giant Swarm

  • Simon Weald

    Site Reliability Engineer, Giant Swarm

– Room 1

  • Attila Szakács-Bertók

    Technical Lead, Axoflow

TBA

– · Room 2

TBA

– · Room 1

TBA

– · Room 2

Closing remarks and raffles

–

Meet the speakers

20 speakers from across the cloud native community. Read their bios and find their talks.

See all speakers
– Room 1 Keynote

Don't panic! Somebody has to pick up the phone at 2AM

  • László Kollár

    Technology Manager, SMP Solutions

AI makes creating software cheaper. Does it make running it cheaper too? Somebody Has to Carry the Pager explores the less-discussed side of AI-assisted development: more software, more dependencies and more automation ultimately reaching the teams responsible for production. From AIOps and AI agents to operational skills, security and accountability, the talk asks a deceptively simple question: when the machine writes and increasingly operates the system, who is responsible when it fails?

About László Kollár

László Kollár is a technology manager with 35+ years of experience in IT transformation, cloud computing, engineering and financial services. He has held senior leadership positions at Microsoft, Morgan Stanley, Cloudera, OTP Bank and SMP Solutions.

Most recently, he built OTP Bank's Cloud Centre of Excellence and led its Azure transformation, before moving to SMP Solutions to develop a standardized cloud platform serving OTP Group.

Throughout his career, László has focused on the intersection of technology, business and organizational change—from building and transforming large engineering and IT service organizations to helping enterprises adopt cloud and emerging technologies. He regularly writes about cloud, AI, technology strategy and their broader economic and organizational consequences on his blog, The Floorshrink Diaries.

– Room 1 Keynote

Escaping the DevOps Hamster Wheel of Busywork

  • Piotr Trębacz

    Senior DevOps Manager, IGT

You automate away millions in costs, yet you remain invisible. You’re the first call when things break, but the last person with the authority to truly fix them. This isn't the strategic role you signed up for - it's a hamster wheel of thankless tasks and shifting priorities, and it's time to get off.

This session delivers a practical framework to break the cycle. No vague advice. Just battle-tested tactics to reclaim your focus and prove your value.

You will learn how to:

- Master the art of strategic pushback. Learn to kill low-impact fire-drills and reframe urgent demands by asking one simple question: "To do this, what gets de-prioritized?"

- Make Your Value Impossible to Ignore with the business outcomes that leaders notice and reward, shifting the narrative from a reactive cost center to an indispensable strategic partner.

- Navigate the Real Power Structure by learn to see the corporate machine as it truly operates, not as it's described. Use this insight to bend the unwritten rules and bypass bureaucracy, becoming an effective rebel whose impact is undeniable.

About Piotr Trębacz

Piotr is a Senior DevOps Manager in the casino industry, championing North Star thinking, Continuous Improvement, and Measurable Outcomes - to transform troubled software and processes into modern standards. He's specialty is combining hard technical skills and soft leadership pragmatism with high agency & Over the past 13 years, he's led $1M+ digital transformations, scaled platform engineering across hundreds of developers and thousands of microservices on AWS/GCP at a time across Air Defense, Gaming, Banking, Logistics, Fuel, Mars Rovers and Gambling & Industry 4.0.

He achieved that with Second Brain of almost 7000 hand-gardened notes bringing unconventional outside-in perspective like his maverick mantras: "Culture Eats Strategy," "Velocity Wins," and "Automation Is Oxygen".

Off-duty, find him in Hawaiian shirts, Game Mastering RPGs, tinkering in his rapid prototyping workshop, or optimizing board game tactics with probability.

Trusted as Code Europe Program Committee member, Ex-GDG Organizer, and speaker, he happily mentors younger experts sharing battle-tested insights.

– Room 1

Every Workload Deserves an Identity Instead of an API Key

  • Zsolt Varga

    CTO, Riptides

  • Nándor Krácser

    CISO, Riptides

Most workloads still reach cloud APIs, databases and SaaS services with long-lived API keys, stored in a Kubernetes Secret or an environment variable and shared by every replica. The key says nothing about which workload is using it, and once it leaks it stays valid until it gets rotated.

AWS, GCP, Azure, OpenAI, Anthropic and other major providers support token federation, which reduces the risk, but the exchanged token can still be leaked or misused, and federating several disjoint systems is far from easy.

In this talk we show how to close that gap by attesting a workload and issuing it a SPIFFE identity, exchanging that identity for short-lived credentials utilizing our open source tokenex library, and injecting them into outgoing requests so the application never holds a secret. We compare this with the usual delivery options (Secrets, mounted files and sidecars) and demo SPIFFE identity based secret injection end to end.

About Zsolt Varga

Zsolt Varga is the CTO of Riptides, a cybersecurity startup focused on workload identity and secure service-to-service communication for AI and cloud-native environments. He has more than 25 years of experience in software development and infrastructure engineering, with a particular focus on cloud-native technologies, Kubernetes and network security. Previously, he was a Senior Software Engineering Leader at Cisco Outshift, joining Cisco through its acquisition of Banzai Cloud.

About Nándor Krácser

Nandor Kracser is Co-founder and CISO of Riptides, where he builds kernel-level workload identity and runtime machine IAM infrastructure using SPIFFE, kTLS, and mTLS. Before Riptides, he was part of Cisco's Emerging Technologies group following Cisco's acquisition of Banzai Cloud, the Budapest-based cloud-native startup where he created Bank-Vaults, the open source project for managing HashiCorp Vault and secrets on Kubernetes. Earlier in his career he worked at IBM and Ustream. His work centers on Linux kernel engineering, eBPF, TLS 1.3, and post-quantum cryptography. He is based in Hungary.

– Room 2

Beyond etcd: Building a Kubernetes-Compatible State API with Kafka

  • Richard Kovacs

    CTO, HariKube

What happens when you use Kubernetes API semantics to expose large-scale application state? This talk explores HariKube, a K8s-compatible API layer powered by distributed databases and Kafka for high-throughput state ingestion and event delivery.

We’ll walk through an end-to-end state transition - from incoming Kafka message to global revision exposure - and unpack key architectural trade-offs:

Preserving K8s API guarantees vs. storage-side scaling

Where Kafka complements K8s watches

Managing limitations like joins, cross-object transactions, and strong consistency

Leave with a practical framework, architectural trade-offs, and a reproducible demo.

About Richard Kovacs

Richard is a Chief Technology Officer at HariKube with many years of DevOps background. His main focuses are Go micro-service, Kubernetes engineering and automation. Richard is passionate about technology and innovation. His constant curiosity drives him to learn, gain knowledge and be an expert in this area. He also loves getting involved with open-source communities. He is also a frequent speaker at local conferences and community events.

– Room 1

The role of platform engineers and DevOps in agentic coding

  • Viktor Nagy

    Fractional Head of Product and R&D, Clearvis PMS Kft

DevOps, Ops, Platform Engineering were always secondary to application engineering. With agentic coding, most of us expects big changes in our processes, and likely an increased valuation of the people who can support their teams to become more efficient by offering them to necessary tools.

In this talk, I'll argue that platform engineers and DevOps people are likely the best positioned or even necessary to implement the company specific integrations needed for reliable, scalable and maintainable agentic coding solutions.

At the same time, to really drive these changes it's not enough to understand the tech and have a vision. I'll share with you the personal and cultural aspects too that you should navigate if you want drive these changes, and we'll discuss briefly the current situation of the industry, as Kent Beck put it: we are back to the times before the Agile manifesto, and we approach our work as they did before 2001.

About Viktor Nagy

At GitLab, I led the development of application delivery features, with a focus on infrastructure as code and Kubernetes integration. My journey in tech started on the engineering side, where I managed production infrastructure across multiple startups before transitioning to product management. This hands-on background gives me a deep appreciation for the challenges DevOps teams face daily and drives my passion for creating tools that make cloud-native development more accessible and efficient.

– Room 2

563 Things Karpenter Didn't Know About

  • Gergely Daroczi

    Project Lead, Spare Cores

  • Attila Nagy

    cloud architect, Spare Cores Kft

Karpenter scales workloads based on provider specs, which often don't reflect real performance. Two instance types with identical vCPU counts can show vastly different single-core speeds, memory throughput, and execution latency -- even at similar CPU frequencies.

In this talk, we bridge real-world compute telemetry into Karpenter scheduling. Using 563 performance benchmarks collected by Spare Cores across thousands of cloud server SKUs, we built a lightweight Kubernetes controller that dynamically updates Karpenter NodePool specs with targeted high-performance instance types, without modifying Karpenter core.

About Gergely Daroczi

Gergely Daroczi, PhD, is a passionate R user and package developer for two decades. With over 15 years in the industry, he has expertise in data science, engineering, cloud infrastructure, and data operations across SaaS, fintech, adtech, and healthtech startups in California and Hungary, focusing on building scalable data platforms. Gergely maintains a dozen open-source R and Python projects and organizes a tech meetup with 1,800 members in Hungary – along with other open-source and data conferences.

About Attila Nagy

Watching the Internet evolve for decades, while designing, building, breaking, fixing, and trying to understand the systems that make it work.

– Room 1

Weights and Measures: What It Takes to Serve an LLM on Kubernetes

  • Márton Enyedi

    Devops Engineer, Origoss Solutions

GPUs bill for existing, not for computing, and that was the first problem in finding out what it takes to serve a large language model on Kubernetes. So the whole stack went up against a simulator first, with no GPU and no model weights. The real hardware lived in a node pool that sits at zero until a pull request scales it up and another scales it back down. Nine of those summon-and-park cycles over two weeks. Then it was running, and slow. Cold start was 26 minutes, 23 of them disk I/O with the GPU idle, and the bottleneck moved every time I fixed one: the volume holding the model, then the image pull, then the boot volume throttling that pull. It ended at nine minutes, four of them the node booting. Serving was still slower than it should have been, so I instrumented the model server and its cache. By then I could measure throughput and latency, but not output quality. That matters, because every lever that makes a model faster can make its answers worse, and the obvious way to measure that trade needs about ninety thousand test questions to resolve an effect this small. The public benchmarks hold a few thousand. This is what it takes, on open source components end to end: the cold-start teardown in order, why the usual alerts for the model's cache stay quiet through the degradation they are meant to catch, and how to find out what a speed-up costs in quality with the questions you already have, once you decide in advance how much worse is still acceptable.

About Márton Enyedi

Márton Enyedi is a Budapest-based DevOps and Platform Engineer working on Kubernetes, automation, and cloud-native infrastructure.

– Room 2

Federating Clusters for Zero-Downtime Kubernetes

  • Dominik Táskai

    Senior Software Engineer, Genesys

When your primary Kubernetes cluster goes dark, how quickly can you restore service? Most organizations rely on backup-and-restore strategies that leave applications unavailable for minutes or hours. Linkerd's federated services flip this paradigm by treating multiple clusters as a unified service fabric.

This session explores architectural patterns that leverage cross-cluster service discovery and intelligent traffic distribution to achieve true zero-downtime resilience. We'll demonstrate live failover scenarios, examine federation deployment models, and dive into the networking mechanics that make seamless disaster recovery possible—all while maintaining security boundaries and operational simplicity.

About Dominik Táskai

Platform engineer specializing in cloud-native and distributed systems, with hands-on ownership of globally distributed platforms on AWS. Active CNCF community contributor, conference speaker, and program committee member. Experienced in incident leadership, observability at scale, and building internal platforms that enable hundreds of engineers.

– Room 1

Temporal as Agent Runtime: Lessons From Self-Hosting It in Production

  • Adam Hegedus

    Engineering Manager, Platform, Zocks

Anthropic suffered an hour-long outage. Our agent workloads recovered without manual intervention, and no engineer replayed a workflow by hand.

This is what durable execution provides. We run our AI agents on Temporal, and we operate Temporal ourselves on Kubernetes.

We build for the financial advice industry. Self-hosting is not a preference in that context. It is how we keep client data, model traffic, and complete workflow history inside a boundary we control and can demonstrate to an auditor.

The session covers both halves of that decision. First, why Temporal suits agent workloads: workflow steps correspond to agent turns, activities to tool and model calls, and signals to human approval. Second, what operating it actually costs. The Temporal services are straightforward; the Cassandra cluster beneath them is not.

Attendees will leave with our Prometheus alerting rules, the KEDA configuration that scales agent workers on task-queue depth rather than CPU utilisation, and a clear account of when a team should choose the managed service instead.

About Adam Hegedus

Adam Hegedus, a tech professional with math background, currently works at Zocks as the Engineering Manager of the Platform team. His expertise spans diverse companies, previously worked in a similar role at Colossyan, and earlier contributing to various open-source softwares like Koperator (operator for Apache Kafka) along with some other Kubernetes-related projects at Banzaicloud (acquired by Cisco).

– Room 2

The Hitchhiker’s Guide to Kubernetes Scheduling and Orchestration

  • Hajnal Máté

    Senior Machine Learning Operations Engineer, Aumovio SE

Kubernetes scheduling looks simple from the outside: a Pod needs a node, the scheduler picks one. But real platforms quickly discover that “where should this workload run?” is not one question. It is many.

Batch jobs need queues and fairness. AI workloads need GPUs and gang scheduling. Multi-tenant platforms need quotas, priorities, preemption, and isolation. Edge or hybrid clusters need placement across failure domains. For agentic workloads everything is about latency. Some teams can stay with the default Kubernetes scheduler, while others eventually reach for Kueue, Volcano, custom scheduler plugins, deschedulers, autoscalers, or external orchestration systems.

In this talk, we will travel through several “worlds” of Kubernetes orchestration and scheduling. For each world, we will look at the problem it solves, the trade-offs it introduces, and when it is probably the wrong tool. The goal is not to crown one scheduler as the answer to everything, but to build a practical mental map for choosing the right orchestration pattern for your workload.

And yes, we will end up with 42 scheduling problems you should ask about before building your next platform.

About Hajnal Máté

Mate Hajnal is a Senior Machine Learning Operations Engineer at Aumovio, following the spin-off of Continental’s automotive branch. With a background in DevOps and Platform Engineering at Red Hat and Nokia, he brings deep expertise in cloud-native infrastructure, Kubernetes, and scalable ML systems. Mate is an active contributor to the CNCF ecosystem and enjoys building bridges between data science, artificial intelligence and modern operational practices.

– Room 1

The Demo Worked. Then Reality Happened: Operating Shared AI Inference on Kubernetes

  • Paul Power

    Software Engineer, Red Hat

Running an AI model on a GPU is straightforward. Turning a few GPUs into shared infrastructure that a team can actually depend on is a platform engineering problem.

This talk follows the evolution of a small three-GPU Kubernetes lab from “we have some spare hardware, let's run some models” into a shared inference platform for developers. The first deployments worked. Then shared use exposed the problems the demo had hidden.

Parallel multi-gigabyte model downloads saturated the network. Limited GPU memory made model placement and concurrency increasingly important. Restarting workloads could mean downloading the same large artifacts again. Adding more GPUs did not solve the surrounding infrastructure constraints. And when something went wrong, limited visibility made it difficult to tell whether the problem was the model, the GPU, Kubernetes, storage or the network.

We'll follow the changes that came out of those failures: controlled model preloading, node-local model caching, GPU-aware workload placement, inference routing, Prometheus and Grafana observability, deployment guardrails, and recovery procedures designed around large model artifacts rather than ordinary container workloads. We'll also look at how the platform is evolving from individual model deployments toward an inference service that developers can use without needing to understand the underlying cluster.

The implementation uses Kubernetes and open-source inference and observability components including vLLM, llm-d, Prometheus and Grafana, but the focus is on the operational patterns and lessons rather than a particular Kubernetes distribution or product.

Finally, we'll look at what happens when a workload outgrows the local GPUs entirely, and how the same deployment and validation patterns can be carried onto temporary larger GPU infrastructure rather than rebuilding the environment from scratch.

This is not a perfect reference architecture. It is a case study of what broke, what we changed, what still has limits, and what we would design differently knowing what we know now.

Attendees will leave with practical approaches for:

controlling the network and storage impact of large model artifacts; making constrained GPU capacity useful across multiple workloads; adding observability that helps distinguish model problems from infrastructure problems; designing deployments and recovery around expensive model state; putting guardrails around shared inference infrastructure; and thinking about AI inference as a platform capability rather than a collection of GPU-backed pods.

The session is aimed at Kubernetes and platform engineers, SREs, DevOps practitioners and developers who are beginning to run AI workloads on their own infrastructure. No deep machine-learning background is required.

About Paul Power

Paul is a software engineer based in Waterford, Ireland. He is currently working on the Q-Fence EU Horizon project, and emerging technologies with OCTO at Red Hat.

– Room 2

Working with filesystem in Time Series database

  • Dmytro Kozlov

    Software engineer, VictoriaMetrics

Time Series databases face the significant challenge of processing vast amounts of data. At VictoriaMetrics, we are actively developing an open-source Time Series database entirely from scratch using Go. Our average installation handles between 2 to 4 million samples per second during ingestion, with larger setups managing over 100 million samples per second on a single cluster. In his presentation, we will explore various techniques essential for constructing write-heavy applications such as: - Understanding and mitigating write amplification. - Implementing instant database snapshots. - Safeguarding against data corruption post power outages. - Evaluating the advantages and disadvantages of utilizing Write Ahead Log. - Enhancing reliability in Network File System (NFS) environments. Throughout the talk, we will illustrate these concepts with real code examples sourced from open-source projects.

About Dmytro Kozlov

Dmytro is a software engineer with experience in scalable applications and enhancing user experiences. With a strong foundation in backend development, cloud systems, etc. Currently working at VictoriaMetrics, focusing on cloud solutions and datasources for VictoriaMetrics and VictoriaLogs. Proficient in languages such as Go, Javascript, TypeScript, he is passionate about leveraging technology to solve real-world problems and improve efficiency.

– Room 1

Managing an Enterprise-Grade AI Infrastructure with More Than 3,000 Nodes

  • Janos Binder

    Staff Engineer, SAP SE

In this talk, we will provide an overview of how we manage a Kubernetes-native environment spanning more than 50 landscapes across six different hyperscalers, including hybrid solutions running on private cloud infrastructure.

We will show how we manage this highly heterogeneous setup, deploy changes quickly and consistently, and abstract away differences between hyperscalers to reduce vendor lock-in. We will also briefly cover how we host open-weight LLMs and proxy access to closed-weight LLMs, as well as the challenges of GPU management and the approaches we use to address them.

Our goal is to share practical experiences and best practices that can help the community tackle the challenges of operating large-scale, multi-cloud AI infrastructure.

***Note to the organizers***: I am also submitting a separate talk on Software Engineering with AI Agent Skills. I will leave it to your consideration which topic is the better fit for the conference. I also had the opportunity to present at KCD Budapest last year.

About Janos Binder

Janos is a seasoned leader with over 15 years of experience in multinational environments, adept at managing global teams and projects. He has collaborated with professionals from more than 50 nationalities and brings deep expertise in developing enterprise software and CI/CD pipelines across various cloud platforms. Holding a PhD in Bioinformatics, Janos balances his technical prowess with a passion for Brazilian jiu-jitsu and quality time with his daughters.

– Room 2

I heard you like virtualization.

  • Gergő Huszty

    Cloud infrastructure engineer, architect, IBM

Turns out, enterprises are still running lots of mission critical applications inside VMs, without any interest in containers. What if they must migrate away from their current platform? Could a public cloud provider create a platform for them, based on Kubernetes, which will meet the traditional requirements of a virtual machine infrastructure? Running Linux with qemu-kvm in a Pod? Fine. What about Windows (2008!) or other, fully proprietary OS instances? With multiple network cards? Migrated from VMware vCenter? Meet Kubevirt, OVN-K, Forklift and others that makes up Openshift Virtualization Engine, from a cloud provider's perspective. The talk shows how to build a service that allows customers to manage their VMs on Kubernetes, in a way that make sense, at scale.

About Gergő Huszty

Gergo spent a decade in the telecommunications industry where he worked on developing carrier grade cloud infrastructures, including Kubernetes. With that he joined IBM to work on IBM Kubernetes Service networking, and now on Openshift Virtualization Service. Gergo had a bunch of talks on Hungarian meetups about both corporate and open-source projects. Once he contributed one character to the Linux kernel. :-)

– Room 1

Stop Breaking Your OTel Collectors: Production Patterns That Actually Work

  • Bence Csati

    Senior Software Engineer, Axoflow

After deploying OTel Collector comes the real-world challenges: filtering noisy telemetry, routing data to multiple destinations, managing resource consumption, and keeping your pipeline reliable as traffic scales. From basic collector deployments to production-grade, this talk shows practical patterns to solve common problems, like:

- Using tail sampling to reduce costs without losing critical traces, - enriching spans with Kubernetes metadata for better debugging, and - implementing batch processing to handle traffic bursts efficiently.

You'll see examples for how to design collector architectures that match your infrastructure (Gateway mode, DaemonSets), configure health checks and self-monitoring, and structure your processor pipelines for performance and maintainability. We'll cover real operational challenges: graceful shutdown to avoid data loss, memory tuning to prevent OOMKills, and strategies for testing collector configs before they hit production.

About Bence Csati

Passionate Software Engineer who loves building reliable systems and actively engages with the cloud-native community to advance Kubernetes security and observability. Maintainer of multiple CNCF sandbox projects, including Bank-Vaults, which is a project dedicated to simplifying the complex world of secret management. The Logging operator, which solves logging-related problems in Kubernetes environments.

Currently, working on the Logging operator and the new Telemetry controller at Axoflow.

– Room 2

From Frankenstein to Kamaji: Lessons in Building a Single CAPI Cluster Across Multiple Providers

  • Xavier Avrillier

    Solutions architect, Giant Swarm

  • Simon Weald

    Site Reliability Engineer, Giant Swarm

At Giant Swarm we use Cluster API to provision and bootstrap our k8s clusters. With this setup, control plane (CP) and worker nodes must run on the same infrastructure which was never an issue so far...

However, in bare-metal environments, using 128-core servers for CP nodes is luxury. It's far more efficient to host them as virtual machines on a hypervisor while keeping workers on physical hardware. But can we get around CAPI's limitations?

We will walk through how we built Frankenstein's cluster by mixing vSphere for the CP and Proxmox for workers as a testing ground. While technically functional, this required "hacky engineering". We will share the hurdles we hit and the operational risks of this hybrid cluster setup.

Finally, we will demonstrate how we solved this challenge with a cleaner, upstream-friendly alternative. Kamaji lets us run the CP as pods in a management cluster. We achieved even better resource optimisation with full native community support and no custom hacks.

About Xavier Avrillier

I work as a Solutions Architect at Giant Swarm, currently working on the managed Kubernetes product in hybrid environments and smart factories. My main focus is around cluster lifecycle and customer implementations.

About Simon Weald

Simon is a Site Reliability Engineer living in the UK, working remotely for Giant Swarm. He’s into keeping infrastructure tidy, running Kubernetes clusters without drama, and helping dev teams ship things reliably.

He spends most of his time thinking about observability, automation, and making sure systems behave when things go sideways. Simon has a strong interest in open source and likes to tinker with new tools in his spare time.

– Room 1

Death of a Log Line: Delivery Guarantees in Telemetry Pipelines

  • Attila Szakács-Bertók

    Technical Lead, Axoflow

It's much easier to detect duplicate logs than lost ones!

Duplicates are loud: you'll deal with them eventually. Lost logs are silent: you might never learn they were dropped, until an alert doesn't fire or a log isn't there when you need it. By then it's too late.

We build telemetry pipelines on a simple assumption: what we send is what arrives. In reality, almost every hop can drop your data without telling anyone, and your real guarantee is set by the weakest link, not the part you trust most. Even OpenTelemetry Collector, by default, treats your data as done the moment it accepts it, fine until it crashes and loses what it already confirmed.

This talk is about that gap: the delivery guarantee you think you have versus the one you really get. We'll see where data quietly disappears (and where it gets duplicated instead), what it takes to prevent it, and why exactly-once delivery is mostly a myth.

Let's stop hoping our logs arrived, and start knowing what our pipeline guarantees!

About Attila Szakács-Bertók

Founding Engineer of Axoflow, leading the dataplane team, specializing in scalable log ingestion and processing pipelines for enterprise and cloud-native environments.

Longtime syslog-ng and AxoSyslog developer with a strong passion for open source and community-driven innovation. Focused on building reliable, high-performance log management and observability systems that address real-world operational challenges.