Work
Senior Software Engineer, contract
Mar 2026 – 20 Sep 2026
I took the contract because it was the kind of ask I like. See a process, and leave the process alone. We shipped that. On 20 September 2026 the contract ended.
Product Reliability Engineer, contract · remote, London entity
Aug 2025 – Mar 2026
Teams were arriving on OpenTelemetry through three parts of the product at once. I stayed on those surfaces, and with the people walking onto them.
- The surfaces were Infra Explorer, the Notification Center, and incidents, SLOs included. A surface becomes real when an account depends on it. That was the bar I kept.
- A lot of the hours were with a customer, halfway through getting OpenTelemetry actually running, and then back with the product, saying what that hour had shown us.
Founding Engineer · remote
Apr 2024 – Jul 2025
We were building the place you open when you need to ask what happened. A tenant could be sending about 15 TB a day, on the order of 170,000 events a second, from the US, Europe, and India. I liked the size. I liked it more when a plain question still received a plain answer.
- Queues were where I dug in. A lag number says something is late. I wanted the produce and the consume as one story, through the messaging conventions, so a queue could be read the way a trace is read. The instrumentation, the lag APIs, and the docs cover Kafka, MSK, Strimzi, and Celery.
- ClickHouse was my favorite room. Tables, views, what deserved an index, how hard we were willing to compress. I was glad to sit there until the store could answer.
- Trace funnels came from a frustration I can still point at. A span was on the screen. How often a request moved from that span to the next one was missing. The first version, then the analytics.
- A call that leaves the process deserves the same care as a call inside it. That became external API monitoring, with the write-up and a later correction to the error rate once the number and the traces had drifted apart.
- I also built a small agent, inside the company, that read the telemetry and offered a possible cause. It stayed inside the company. The rest of the time went to scope, to customers, to the cluster, and to the pager. I have a lot of respect for a quiet pager.
Software Engineer, Platform Observability
Dec 2023 – Mar 2024 · Amsterdam
Four months in Amsterdam, inside payments. I arrived holding one picture: a single pipe, where the trace, the metric, and the log obviously belonged to each other.
- The pipe was the OpenTelemetry Collector. Two things mattered past the demo. Another engineer had to be able to operate it a year later. And after a rough day, the three signals still had to agree.
- The parts that grew with traffic moved onto KEDA. Personal data came off the traces as they passed through.
- We put load on it before we asked anyone to call it production. That order matters to me.
Jul 2020 – Aug 2023 · Bengaluru
Three years in Bengaluru. Some of it was the pipe that took the logs in. Some of it was the mesh between clusters. Under both, I got stubborn about a small decency in Kubernetes. When a log speaks, it should name the object it is speaking about.
- I wrote a collector in Go. It batched, compressed, and sent about 5 million OTLP logs an hour, with the pod already named on the line. The person on call already has the name. They can spend the night on the actual fault.
- The shared ingestion path was processors on the collector, around 20 million events an hour. The figure I still tell people is the modest one. 1 vCPU and 1 GiB was enough.
- Across clusters, traffic should find its way when a zone goes away. Locality-aware balancing, and one control plane over the Istiod fleet. Those were the same years I spent on Kubernetes logging, KEP 1602 and KEP 3077, so a line could carry its object and its context together.
- I also took the Java agent into Play, WSO2, Cassandra, and Tomcat. Agent work is quiet. It unblocked a lot of teams, and I was glad to be the one doing it.
OpenMF looks at Android phones after the fact, for a forensics case. I proposed in 2019 and missed. I proposed again in 2020 and got in. Writing the second proposal is one of the better decisions I have made.
- The 2020 proposal is the one that got me in. I built the way through a case: contacts, calls, messages, and a report at the end that a person could read. The write-up is in the SCoRe Lab archive.
- Clio, in 2019, was a small system for third-party licences. It was not taken. I kept the proposal. The idea was decent, and the miss taught me how to write the next one.
- In 2021 I came back as a mentor, for three students. Swapnal Shahil wrote about week six, and I am in that note. I liked their bugs almost as much as my own.
Software Engineer, intern
May 2019 – Jul 2019 · Bengaluru
One summer on a lock scheduler. What has stayed with me is the week the schedule grew large and the scheduler forgot its manners.
- I wrote it and hung it off SmartThings Cloud. Node.js, OAuth, DynamoDB. The wiring was ordinary. The question was whether a lock still meant what it said once everyone arrived at once.
- I drove the concurrency up to around 7 million operations and repaired what gave way. A short internship. It taught me to keep testing past the point where a system is still being polite.