Distributed Systems · June, 2025 — Present
Software Engineer, Atlan
At Atlan, I work on backend and platform systems that sit behind a large-scale metadata platform. My work spans distributed processing, workflow infrastructure, event-driven systems, and internal engineering tools. A recurring part of the job is dealing with systems that work well under normal conditions but become interesting when they have to handle scale, failures, or unexpected states. I've gradually moved from solving individual production problems to designing the systems and abstractions that prevent those problems from recurring.
What I worked on
- Building and improving distributed backend systems where reliability, consistency, and failure recovery matter.
- Designing infrastructure for long-running workflows and asynchronous processing, with a focus on making failures recoverable rather than disruptive.
- Working on reusable systems for metadata lineage, real-time state propagation, and engineering automation.
- Building practical AI tooling that can interact with real engineering environments, including code execution, Kubernetes operations, debugging, and automated workflows.
- Taking ownership of problems across service boundaries, from understanding the root cause to designing, implementing, and shipping the fix.
Technologies
Python · Kafka · Numaflow · Kubernetes · AI Agents · Distributed Systems