Skip to content
Go back

How Instagram Runs Smoothly: Multi-Process Architecture and C++ NDK in Android

Instagram Multi-Process Architecture

Instagram constantly streams videos, decodes bitmaps, and applies image filters — yet the feed scrolls smoothly. If we put all of that work inside a single process the way a typical app does, the garbage collector would fire every few milliseconds and the UI thread would stutter. So what does Instagram actually do differently?


The Root Problem: One Process, One GC

In a standard Android app, everything runs inside a single process. That process has one JVM heap — and one garbage collector watching over it.

Every time a large video chunk is downloaded and decoded, memory is allocated and then released. Frequent allocation/deallocation is one of the fastest ways to trigger GC. When GC runs on the main (UI) thread, the screen cannot refresh until it finishes. Miss the 16 ms frame window and the user sees a dropped frame — that choppy feeling we call lag.

The bigger the heap, the longer each GC run takes:

Heap SizeApprox. GC TimeTask (10 ms) + GCResult
50 MB~2 ms12 msSmooth ✅
200 MB~8 ms18 msLag ❌

The naive fix — android:largeHeap="true" — makes this worse by giving GC more memory to scan on every run.


Solution 1: Split Work Across Multiple Processes

Android allows a single app to spawn multiple processes. Each process has its own isolated memory space and, crucially, its own independent garbage collector.

Instagram’s architecture roughly looks like this:

When GC fires repeatedly inside a worker process because of a large video download, the main process never knows it happened. The UI thread keeps drawing frames without interruption.

Processes talk to each other through Inter-Process Communication (IPC). Each process also has its own separate GC instance — they do not share a garbage collector.

Key rule: Anything that causes frequent memory churn should live in a process that does not own the UI thread.


Solution 2: Native Memory via C++ NDK

Even with multiple processes, the JVM heap has a ceiling. Any object created with new in Kotlin or Java counts against that limit and is subject to GC.

C++ code accessed through the Android NDK (Native Development Kit) allocates memory directly from RAM using the OS, completely bypassing the JVM heap. This memory is not counted against the app’s heap limit and is never touched by the Java garbage collector.

Instagram (and its parent library, Fresco) stores decoded bitmaps in native memory for exactly this reason. Compare the two approaches:

LibraryLanguageBitmap StorageGC Impact
FrescoC++ (via JNI)Native heapNone — GC never sees it
GlideJava/KotlinJVM heapGC scans and may pause for it

The bridge between Java/Kotlin code and C++ is called JNI (Java Native Interface). Your Kotlin code calls a native method, and the JNI layer routes that call to the corresponding C++ function. The C++ code can then allocate, manipulate, and free memory on the native heap without involving the JVM at all.


Why C++ Is Faster for Compute-Heavy Tasks

GPU Parallelism

When Instagram applies a filter that involves matrix operations, C++ can distribute those calculations across multiple GPU cores simultaneously. Java executes the same loop sequentially — one element at a time. For a 3×3 matrix multiplication:

This is why compute-heavy operations (filters, video decoding, ML inference) are almost always implemented in native code on Android.

APK Compression: A Real-World Example

Suppose an app needs 100 MB of offline data bundled with it.

The C++ solution:

  1. Compress the 100 MB data to ~4 MB using a native compression library.
  2. Ship a ~10 MB APK (base APK + 4 MB compressed asset).
  3. On first launch, decompress in C++ — fast enough that the user barely notices.

Java-based decompression of 100 MB is slow because it runs on a single thread and goes through the JVM. C++ decompression handles the same data significantly faster by operating closer to the hardware and, where supported, using SIMD instructions.


Thread Count and Memory Cost

One more thing the class notes touch on: threads are not free.

Each thread requires its own stack. On Android, a thread stack typically occupies between 0.5 MB and 1 MB depending on the device. Creating 1,000 threads would consume 500 MB to 1 GB of memory just for stacks — before those threads do any actual work.

This is why Android apps use thread pools (via ExecutorService, coroutines with Dispatchers.IO, etc.) rather than spawning raw threads for every task.


Summary

TechniqueWhat It Solves
Multiple processesIsolates heavy GC from the UI thread
C++ NDK / native memoryAllocates outside the JVM heap, no GC involvement
JNI bridgeLets Kotlin/Java call C++ code seamlessly
Thread poolsAvoids ballooning memory from uncapped thread creation

A smooth Instagram feed is not magic — it is a deliberate architectural decision to keep the UI process lean while offloading everything memory-intensive to worker processes and native code.


Share this post on:

Previous Post
Android IPC: When to Use Content Provider vs AIDL
Next Post
How to Configure FreeSWITCH ESL for External Access