Skip to content

On February 19, 2014, Facebook announced that it would acquire WhatsApp. Facebook’s press release described a service used by more than 450 million people each month, with 70 percent of them active on a given day, and it said that WhatsApp was adding more than one million new registered users every day. The press release did not say how many people worked at WhatsApp. That number came from other places. On the same day, Sequoia Capital, WhatsApp’s lead investor, published a post by Jim Goetz that read, “With only 32 engineers, one WhatsApp developer supports 14 million active users, a ratio unheard of in the industry.” TechCrunch reported that WhatsApp employed around 50 people in total. Sixteen days later, on March 7, a WhatsApp engineer named Rick Reed gave a talk at the Erlang Factory conference in San Francisco with the title “That’s Billion with a B.” His second slide described his team in two short lines: “Small (~10 on Erlang)” and “Handle development and ops.” So the same company was described that month as 32 engineers, as about 50 people, and as roughly ten people who both built and operated the server software.

These figures do not agree, and it is important to understand why. Sequoia was a direct beneficiary of the sale, it published on the day the deal was announced, and it gave no method for counting its 32 engineers. TechCrunch reported a figure for all employees. Reed counted the group that worked in Erlang, the language and runtime that WhatsApp’s servers ran on. None of the three sources explains exactly who was included. But all three agree on the essential point: the service was enormous, and the group of people who built and ran it was very small. That combination is the case this piece examines, and it introduces a concept I will refer to as systems leverage. The idea is that WhatsApp’s small team was not the starting point of its success. It was the output of systems work. A few engineers chose a runtime built for very large numbers of simultaneous connections, then measured, tuned and patched that system for years, and the headcount stayed small because of that work. Reed’s own summary of that work fits on one line: decouple, parallelize, patch, monitor.

So how could so few people run a service of that size? How can a single computer hold two million conversations open at the same time? What did the Erlang runtime and the FreeBSD operating system give the team, and what did the team still have to build for itself? What method did Reed’s group follow as the service grew from one server to hundreds? What went wrong along the way? And what does any of this teach us in 2026, when a company with two employees reports $401 million in annual revenue and a founder says he did not write one line of his product’s code?

What Is an Open Connection?

A messaging service must deliver a message to a phone at the moment it arrives. One way to do this is for every phone to ask the server, every few seconds, whether anything new is waiting. With hundreds of millions of phones, that would mean billions of questions, and almost all of them would receive the answer “nothing.” The alternative is for each phone to keep one connection open to a server, so that the server can push a message down that connection whenever there is one to deliver. Each open connection is a socket: a small record that the operating system keeps for as long as the connection lives, together with some memory for the data waiting to be sent or read, and whatever state the server software holds for that user. The important property of these connections is that most of them are idle most of the time. Reed’s 2014 deck lists 147 million concurrent connections and a peak of 342,000 incoming messages per second. Set side by side, those two figures mean that even at the peak, fewer than one connection in 400 would carry an incoming message in any given second. The job of a chat server, then, is mostly to hold a very large number of quiet connections at low cost, and to act quickly on the few that wake up. Capacity becomes a question of cost per connection: how much memory each one needs, how much bookkeeping the operating system does for it, and how evenly the processors share the work when connections become active.

Two Million Connections on One Server

On January 6, 2012, WhatsApp’s engineering blog published a short post titled “1 million is so 2011.” It reported that the team was “now able to easily push our systems to over 2 million tcp connections,” and it gave the measured figure: 2,277,845 open sockets on a single server. The machine had 24 cores (Intel Xeon X5675 at 3.07GHz) and 96 GB of memory, and it ran FreeBSD 8.2 and Erlang R14B03. With all of those connections open, the processors were still 41.9 percent idle.

We can work out what each connection cost. If we divide 96 GB of memory by 2,277,845 connections, we get a little over 40 kilobytes per connection. That figure is a ceiling, not a measured cost. The runtime’s own heaps, the in-memory tables, and the code all share the same memory, so the real budget for one socket, its buffers and its state was smaller than that. The post does not give a per-connection figure, and neither deck does. The post ended with a hiring call, so WhatsApp had a reason to publish an impressive number.

When Reed presented at Erlang Factory on March 30, 2012, his slides told the longer version of the same story. The team had started at about 200,000 connections per server and set a target of one million. It reached 2.8 million, fourteen times the starting point and nearly three times the target, on a machine with 24 logical CPUs and 100 GB of memory. One slide gives the only reason for the runtime in the whole deck: “Erlang has awesome SMP scalability: >85% cpu utilization across 24 logical cpus.” SMP, or symmetric multiprocessing, means that one program uses many processors at once. A runtime that keeps all 24 of them more than 85 percent busy is spreading the work evenly, rather than letting it pile up on a few processors while the rest wait. That is what the runtime gave the team. But the same deck shows how much the team did on top of it. They took a newer kernel timecounter and a newer network driver from a later FreeBSD and applied them to the version they ran, and they tuned the operating system’s settings. Neither deck explains why WhatsApp chose Erlang or FreeBSD. The 2014 deck says only “Awesome choice for WhatsApp — Scalability — Non-stop operations,” and both decks spend their slides on the work needed to make the choice hold.

WhatsApp in 2014: The Numbers on Reed’s Slides

Two years later, the scale had changed completely. Reed’s fourth slide in March 2014, titled simply “Numbers,” lists 465 million monthly users, 19 billion messages in and 40 billion out per day, 147 million concurrent connections, 230,000 peak logins per second, and 342,000 peak messages in per second with 712,000 out. His hardware slide lists about 550 servers plus standby gear, of which about 150 were chat servers holding about one million phones each. The whole platform ran on more than 11,000 cores, on “FreeBSD 9.2 — Erlang R16B01 (+patches).”

We can see how these figures fit together. About 150 chat servers at about a million phones each gives about 150 million connections, which matches the 147 million concurrent connections on the numbers slide. In other words, the million-connection target that was a benchmark in 2012 had become the ordinary operating level of a typical chat server in 2014. And by Reed’s own count, about ten engineers built and operated the Erlang system that carried them. It is tempting to divide 147 million connections by ten and set the result beside Sequoia’s ratio of one developer to 14 million users. We should not. The two ratios measure different things: connections per Erlang engineer on one side, monthly users per an undefined count of 32 engineers on the other. The honest statement is smaller and still remarkable. A group of about ten people ran the server software for a service of this size.

Reed’s Method: Decouple, Parallelize, Partition, Patch, Monitor

The middle of the 2014 deck is not about the language. It is about a method, and Reed summarized it on one slide as a sequence: “Decouple, Parallelize, Decouple, Optimize/Patch, Decouple, Monitor/Measure, Decouple.” The word “decouple” appears four times, and that repetition shows us how central it was to the method. Decoupling means arranging a system so that one part does not have to wait for another. For example, the deck shows the team avoiding transactions in mnesia, the database they ran inside Erlang, because a transaction ties operations together until they all complete, and using a looser mode called async_dirty instead. The team also used separate queues for reads, for writes, and for traffic between machines, so that a backlog in one would not stall the others. Parallelizing and partitioning mean splitting the work of a service into pieces that can run side by side. Each service in the deck was split into between 2 and 32 partitions, addressed through Erlang’s pg2 process groups, so that the load, and any problem, fell on a fraction of the system rather than on all of it. Optimizing and patching went below the application, into the runtime itself. In plain terms, the team changed how the engine beneath their own code allocated memory and scheduled work. The deck lists custom patches to the Erlang virtual machine, known as BEAM: changes to memory allocator configuration, real-time scheduler priority from the operating system, larger buffers for traffic between machines, and round-robin scheduling of asynchronous file input and output. Finally, monitoring meant that every node reported its measurements every second, and the results were pushed to Graphite, a tool for storing and graphing those measurements. None of these steps is exotic on its own. What made them work was that a few people understood the whole system well enough to decide which step each problem needed.

An Example: The Account Table

One incident in the deck shows the method at work, and we can understand it best by starting with how a hash table finds a record. A hash table has a fixed number of slots, called buckets. When a record is stored, a hash function turns its key into a bucket number, and the record goes into that bucket. If the function spreads the keys evenly, each bucket holds only a few records, and finding one means jumping straight to the right bucket and checking a handful of entries. The records that share a bucket form a chain, and the length of the chain is the cost of every lookup. WhatsApp kept its account data in a mnesia table split into 512 fragments, and the in-memory hash tables underneath had a target chain length of 7. After the team added more hosts, the account table became slow. When they looked, the chains had grown past 2,000 entries. Reed’s slide records the moment in a few words: “Hash chains >2K (target is 7). Oops.” A lookup that should have checked about 7 entries could now walk through hundreds or thousands of them, close to 300 times the work in the longest chains. The fix was a patch that changed the hash seed, the starting value mixed into the hash function. With a different seed, the same keys scatter differently across the buckets, and the fix restored the table’s performance. There are three lessons in this small story. First, the problem appeared as a number, a chain length hundreds of times its target, and only a team that measured its system closely would see it. Second, the cause sat below the team’s own code, inside the hash tables the runtime provided. Third, the fix required people who were willing to change how the runtime behaved.

Patches, an Outage, and a Growing Team

Reed’s deck also records what this approach cost. Slides 23 through 27 are spent on the custom changes to BEAM that we saw above. The team was maintaining its own modified version of the runtime, and that is real engineering effort, carried by the same group of about ten people who wrote the application and ran the servers. The method also did not always succeed, and Reed said so on a slide titled “Clearing the minefield.” The slide calls the failure the “2/22 outage.” It began with a glitch in a back-end router, which was followed by a mass disconnection and reconnection of nodes, and that resulted in what the slide calls “a novel unstable state.” The team could not stabilize the cluster, especially pg2, the same process-group mechanism that the partitioning scheme relied on. The outage ended with a “full stop & restart,” which the slide notes was the “first time in years.” Finally, the team did not stay that size. In April 2016, Yahoo Finance reported that Reed, speaking at Facebook’s F8 conference, said “WhatsApp now has 57 engineers, total,” for a service with over 900 million users. That figure has the same limitation as Sequoia’s, because no definition of “engineer” was given. But if we take both figures as reported, the team roughly doubled while the user base roughly doubled, so the ratio of users to engineers stayed close to where it had been. Taken together, these facts show us what systems leverage costs: constant measurement, a willingness to maintain the runtime itself, and the occasional failure that no method prevents.

The One-Person Billion-Dollar Company

The idea has a clear origin. In February 2024, Fortune reported that Sam Altman, who runs OpenAI, told the Reddit co-founder Alexis Ohanian about a betting pool among his friends for the first year of “a one-person billion-dollar company. Which would have been unimaginable without AI and now will happen.” It was a prediction from an interested party, with no method attached. In 2026 the idea has cases, and two are worth setting beside Reed’s slides. The first is Medvi. On April 6, 2026, NewsNation reported that Medvi had two employees, Matthew Gallagher and his brother Elliot, and $401 million in revenue in 2025. Gallagher said, “It’s not an AI company, but I did it with A.I.” The same report names two partner companies, CareValidate and OpenLoop, and the medical and pharmacy work runs through them. At WhatsApp, Reed’s team did not write Erlang or FreeBSD, and the “~10 on Erlang” on his slide did not count the people who did. Medvi goes a step further. Its two employees sit on top of systems that other companies operate, and a headcount of two does not count the people who run them. The revenue figure is the founder’s own, with no audit named.

The second case teaches a different lesson. Moltbook is a platform where, by the count of the cloud security company Wiz, AI agents outnumbered human accounts 88 to 1. Its founder, Matt Schlicht, wrote on X on January 30, 2026, “I didn’t write one line of code for @moltbook. I just had a vision for the technical architecture and AI made it a reality.” On February 2, 2026, Wiz reported that Moltbook’s database was open. Moltbook kept its data in Supabase, and the key that connects to that database sat in the JavaScript that each visitor’s browser downloads, so any visitor could read it. A key in that position does not protect anything by itself. The protection must come from rules inside the database, called Row Level Security policies, which decide which rows each user may read or change. For example, one rule might let a user read only the rows that belong to that user. Moltbook’s database had no such policies, so the key gave “full database access to anyone who has it.” The exposed data included 1.5 million API keys and 35,000 email addresses. According to Wiz, the team fixed it within hours.

We can now set Moltbook beside the account table in Reed’s deck. WhatsApp’s team saw a number drift from its target of 7 to more than 2,000, traced the cause below its own code, and patched the runtime. Moltbook shipped without a basic access rule on its database. Both teams acted quickly once they saw the problem. What separates the two cases is where the judgment came from. At WhatsApp, the people who wrote the code also ran the system, measured it every second, and knew which layer each number came from. At Moltbook, the code arrived, and the judgment about the running system did not arrive with it. Together, the two cases teach one lesson. A headcount counts the people on the payroll. It does not count the systems, partners and judgment that the work depends on.

Toward Small Teams That Understand Their Systems

We can now return to the group of about ten people on Reed’s second slide. The same slide says that Reed joined WhatsApp in 2011 and “Learned Erlang at WhatsApp.” The judgment that made the team small was learned on the job, on a live system, by people responsible for both development and operations. I believe this is the right way to read the small companies of 2026. By their founders’ own accounts, AI helped two people build a business with $401 million in revenue and helped a founder build a platform without writing its code, much as Erlang and FreeBSD let about ten people carry 147 million connections. I am confident that small teams will keep doing things that once required large ones. But the WhatsApp story shows that the leverage came from knowing what to decouple, where to partition, and when to patch the runtime, and Moltbook shows that this knowledge does not arrive with the code. So we have an obligation, whether we lead a team or work in one. Before we accept that AI has made the team unnecessary, we must learn our systems the way Reed’s group learned theirs: measure them every second, understand the layer beneath our own code, and put the basic rules of the running system in place before the first user arrives.

Next: Your AI agents don’t scale like web servers

Read next