技巧

Peter Mattis谈分布式数据库与AI编程

精选理由

Peter Mattis回顾了GIMP、Gmail和CockroachDB的开发历程,分享了分布式数据库和AI编程的实战经验。

Peter Mattis分享了从GIMP到Gmail再到CockroachDB的技术历程。他优化了Google的std::map和Go的map数据结构,性能超越标准库。在访谈中,他探讨了分布式存储瓶颈、一致性模型和Raft共识算法。AI工具帮助他重新投入编程,并改进了代码质量。

原文 · Gergely Orosz

Wow: @petermattis has quite the history in shipping: co-created GIMP, built Gmail’s original backend, shipped Google’s Colossus, and is the cofounder and CTO at Cockroach Labs. We talked distributed DBs:

Timestamps:

00:00 Intro 02:42 Peter’s path into tech 04:00 Building GIMP 09:30 Working on Gmail at Google 14:51 Google’s infra: google3, build files, Bazel, and Colossus 21:30 Distributed storage bottlenecks 23:59 Latency, throughput, and availability 30:04 Contributing to libraries 41:52 Google Spanner 46:10 CockroachDB 52:00 Manual vs. automatic sharding 55:28 Consistency models and strong consistency 1:00:03 Raft consensus 1:06:15 How AI brought Peter back to coding 1:19:12 Peter’s tools and agentic workflows 1:23:08 How AI can improve quality 1:26:39 Code reviews: are they done? 1:29:17 100x engineers 1:35:33 Peter’s advice for leveling up your engineering skills

Brought to you by: • @turbopuffer – the team are completely redesigning their storage architecture from first principles. And they’re documenting all of it! Follow along,: https://t.co/h4kBELKz20

• @linear – most of us work with agents in a “single-player” setting. Linear’s take is agent work should be teamwork, and they built it so: https://t.co/knIF2mcz9p

• @WorkOS – fresh from the WorkOS oven: Airlock, the authorization layer for AI agents. It evaluates every request against the agent’s intent and your rules and either allows it, denies it, or routes it to a human for approval. https://t.co/OzuIJ1wHtZ

Three fascinating parts from this convo:

1. Peter was recruited to Google because of... GIMP!

The very first version of the famous Google logo was made in GIMP, the free, open-source raster graphics and image editing software created by Peter and his college roommate Spencer Kimball.

In 2001, Sergey reached out to Peter, invited him to an interview and then made an offer. While Peter said no back then (thanks to the commute), he joined a year later, in 2002.

2. Peter twice “beat” standard library data structures in performance terms.

At Google, one of his colleagues noticed that std::map showed up in memory profiles. Peter looked closer and figured out that the std::map implementation is a red-black tree, meaning every node has two pointers. Peter built a B-tree with nearly the same semantics which was both faster, due to spatial locality, and smaller because it used fewer pointers. They used this data structure inside Google.

Peter also did something similar with Go’s map: it was performant, but he built a Swiss Table implementation that was faster. That implementation later made it into the Go library, with the Go team helping to finish it!

3. If you squint hard, everything in distributed databases and storage systems starts looking like a B-tree.

These B-trees are a recurring theme in this podcast episode: the backend of Gmail, the std::map replacement, CockroachDB’s range index, etc. There’s even a paper on this phenomenon, The Ubiquitous B-Tree