Skip to content
José da Cruz

Enterprise Architect, and Life Thinker

José da Cruz

Enterprise Architect, and Life Thinker

How Instagram-Style Feeds Work: Fan-Out on Write vs. Fan-Out on Read

josedacruz, August 2, 2026August 2, 2026

Imagine you’re on the backend team at a small photo-sharing app called Snapstream. Every user has a home feed: a scrolling list of photos from the people they follow, newest first. For your first 10,000 users, this is easy. For your next 10 million, it turns into one of the trickiest problems in system design. This post walks through why, using the two main approaches real companies use to build a feed, and the surprising problem that breaks the simplest one.

What a feed actually has to do

A feed sounds simple: show me the newest posts from everyone I follow, sorted by time. But think about what has to happen behind the scenes every time someone opens the app. The system has to look at every account you follow, find their recent posts, mix all of those posts together, sort them by time, and hand you the top 20 or so. If you follow 300 people, that’s 300 separate “what did they post?” lookups happening in the blink of an eye, for every single person who opens the app, every single time they refresh.

There are two fundamentally different ways to solve this, and the difference comes down to a simple question: do you do the work when someone posts, or when someone reads?

Option 1: Fan-out on write (the “push” model)

The first approach is called fan-out on write. “Fan-out” just means one thing spreading out to touch many things — like a single deck of cards fanning out into many separate cards. Here, the “one thing” is a new post, and the “many things” are the inboxes of everyone who follows the poster.

Here’s how it works. The moment you post a photo on Snapstream, the system doesn’t wait around. It immediately looks up your entire list of followers and copies a reference to your new post into each of their personal, pre-built feed lists — think of it as a mailbox that belongs to each follower. If you have 500 followers, posting triggers 500 tiny writes, one per follower’s mailbox, usually handled by a background job so you don’t have to wait for all 500 to finish before your “Post” button confirms.

The payoff comes later, when any of those followers opens their app. Their feed is already sitting there, pre-assembled, sorted, and ready. Reading it is one fast lookup — pull the mailbox, done. This is why the model is popular: reads are the thing that happens constantly (every time anyone opens the app), while writes (posts) happen far less often. Fan-out on write moves the expensive work to the rare event and makes the frequent event cheap. That’s usually a great trade.

The celebrity problem

This is where it gets interesting. Fan-out on write works beautifully until someone with 50 million followers posts a photo. Now that one post has to be copied into 50 million separate mailboxes. Even if each individual write is fast, 50 million of them takes real time and real infrastructure — and it can back up the whole background job queue behind it, delaying everyone else’s much smaller fan-outs too.

This is known in the industry as the celebrity problem, and it’s not hypothetical — it’s the reason Twitter (now X) famously had to rethink its own feed architecture as it grew. A design that’s perfectly efficient for a normal user becomes a bottleneck the moment a small number of accounts have an enormous audience.

Diagram comparing fan-out on write, where a new post is pushed into every follower's inbox immediately, versus fan-out on read, where posts are stored once and merged together at the moment a user requests their feed
Fan-out on write pushes work to post time; fan-out on read defers it to request time.

Option 2: Fan-out on read (the “pull” model)

The second approach flips the trade-off entirely. In fan-out on read, posting is nearly free — you just save the post once, in one place. Nothing gets copied anywhere. The work happens later, at read time: when you open your feed, the system looks at everyone you follow, fetches their most recent posts on the spot, merges them together, sorts by time, and returns the result to you.

This solves the celebrity problem instantly, because a celebrity’s post is written exactly once no matter how many followers they have. But it reintroduces the original pain: if you follow 300 accounts, opening your feed now means 300 lookups happening live, every time. For an app with millions of daily active users constantly refreshing their feeds, that’s an enormous amount of repeated, expensive work — arguably worse than the celebrity problem, just distributed differently.

What real systems actually do: a hybrid

Neither pure approach survives contact with a large, realistic user base, so production systems typically combine both, splitting users into two groups based on follower count.

  • Normal accounts (the vast majority of users) use fan-out on write. Their posts get pushed into followers’ feeds immediately, because their follower counts are small enough that this is cheap.
  • High-follower accounts (celebrities, large brands, viral accounts) are excluded from the push process entirely. Their posts are just stored, not fanned out.
  • At read time, your feed is assembled from your pre-built mailbox (fast, covers most of what you follow) merged with a live lookup of just the small number of celebrity accounts you follow (slower, but there are only a handful of those per user, not hundreds).

This hybrid gets the best of both: cheap reads for the common case, and no runaway fan-out job when a huge account posts. The cost is complexity — you now need logic to classify accounts, a threshold for what counts as “high-follower,” and a merge step at read time instead of a single simple lookup.

Why this matters even if you’ll never build a feed

Most engineers will never build the next Instagram. But the underlying question — do we pay a cost at write time, or defer that cost to read time? — shows up constantly in smaller systems too. A reporting dashboard can pre-compute nightly summaries (write-time cost, fast reads) or calculate everything live when someone opens it (no write-time cost, slower reads). A search index can update the moment data changes, or scan on demand. The names change, but it’s the same fork in the road.

The heuristic worth remembering is this: push work to whichever side of the read/write equation happens less often, and watch out for the rare event that’s disproportionately large. A “small” percentage of celebrity accounts or nightly reports can quietly dominate your system’s cost if you don’t design for that skew up front. The best architectures usually aren’t a single clean answer — they’re a plain default with a deliberate exception carved out for the outliers.

Related

architecture architecturecase studydistributed systemsscalabilitysystem design

Post navigation

Previous post
Next post

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

  • architecture (154)
  • artificial-intelligence (6)
  • books (2)
  • courses (4)
  • decision-frameworks (4)
  • finances (1)
  • java (8)
  • life-improvement (20)
  • metrics (3)
  • observability (2)
  • puzzles and challenges (3)
  • reviews (1)
  • security (8)
  • spring-boot (36)
  • Systems Thinking (3)
  • GitHub Weekly: MCP Tools Trending Right Now (September 28 – October 4, 2026)
  • The Bulkhead Pattern: Why One Slow Dependency Shouldn’t Sink the Whole Ship
  • Essential Reading: A Philosophy of Software Design — Fighting Complexity One Module at a Time
  • GitHub Weekly: Top 10 Trending Repos (September 21 – September 27, 2026)
  • The Volunteer Tax: Why Raising Your Hand in a Meeting Makes You the Permanent Owner
©2026 José da Cruz | WordPress Theme by SuperbThemes