🧵[1/20]
Welcome back to Cryptolytic! Today we will explore @WalrusProtocol, the decentralized blob storage protocol built on @SuiNetwork, and what advantages it has compared to other decentralized storage designs, especially on recovery and self healing cost.
1
27
🧵[2/20]
I have not posted in ages 😅 I have been buried in research, slides, and Twitter Live lectures on @contributedao.
I have already covered Walrus and more in a recorded session. You can watch it on the @contributedao page.
1
18
🧵[3/20]
I did not write a Substack version because this one has a lot of math and equations. My explanation may not be rigorous, but I promise it is very clear. Substack is painful for equations, so I wrote it in Overleaf instead. You can read it and download the PDF here:
overleaf.com/read/tzqrbbwqwp…
1
16
🧵[4/20]
Big picture, decentralized storage usually comes in two approaches
1.Full replication
2.Erasure coding
Walrus uses erasure coding, but with a twist. Self healing is much cheaper.
WTH is self healing cost? Keep reading.
1
9
🧵[5/20]
Full replication
Easy. Network has 100 nodes, so you store 100 copies of the data across nodes.
But it is insanely expensive.
1GB of data means the network stores 100GB total.
1
8
🧵[6/20]
Just split it
You might say, ok fine, split 1GB into 100 chunks and give each node 1 chunk.
Problem: If some nodes go down or act maliciously, you cannot reconstruct the original file. One missing chunk and you are wrecked.
1
10
🧵[7/20]
Erasure coding
Erasure coding adds redundancy so you can recover data even if some pieces are missing.
Toy idea.: Expand data from size B to 3B.
As long as you collect B out of 3B, you can reconstruct the original.
1
9
🧵[8/20]
This is why DA and storage systems love erasure coding. You cut storage overhead while keeping classic BFT style safety thresholds.
We will use the usual model. The network has 3f+1 nodes, where f is the maximum faulty nodes.
1
9
🧵[9/20]
1D erasure code toy model
Network has 3f+1 nodes.
Split the data into f+1 chunks, then encode it into 3f+1 chunks by adding redundancy, then store one chunk per node.
Total storage overhead is approximately 3x for large f.
Example. N equals 100, so f equals 33. Split into 34 chunks, encode into 100 chunks, and each node stores 1 chunk.
1
11
🧵[10/20]
So far so good. Now suppose a node crashes and loses its chunk.
No big deal, right? The network still has enough chunks to reconstruct.
Here is the problem:
To rebuild one lost chunk, the naive repair is basically this. Download f+1 chunks, which is roughly the full original file size, reconstruct everything, re-encode, then keep only the missing piece. So a node stores about 3/N of the file, but to heal itself it downloads about 1 full file. That is about N/3 times more bandwidth than what it lost.
1
6
🧵[11/20]
Concrete example:
N equals 100, f equals 33, file size equals 1GB. Split into 34 chunks, so each chunk is about 1/34 GB. Encode into 100 chunks, so the total encoded size is about 3GB. Each node stores 1 chunk, about 1/34 GB.
If one node loses its chunk, it downloads 34 chunks, about 1GB, just to recover about 1/34 GB.
That is around 34x more than needed. WTH?
1
6
🧵[12/20]
Walrus fixes this with a 2D erasure coding protocol called Red Stuff, inspired by the Twin code framework.
Tradeoff: Each node stores about 4.5/N of the file instead of 3/N in the toy 1D case, which is 1.5x more.
Key advantage: Recovery bandwidth becomes proportional to the lost data, not the full file. Meaning if the original data is 1GB and a node stores 1MB, then if it loses that 1MB, it only needs to download about 1MB to recover it, not the full 1GB.
1
7
🧵[13/20]
How it works at a high level. Still assume the network has 3f+1 nodes.
Split the original file into a 2D grid with f+1 rows and 2f+1 columns.
Then encode in two dimensions.
Primary encoding extends columns.
Secondary encoding extends rows, using a copy or second view.
Result. Every node stores a pair. One primary sliver and one secondary sliver.
1
10
🧵[14/20]
Primary encoding, column wise
You have 2f+1 columns, and each column has f+1 chunks.
Encode each column from f+1 chunks into 3f+1 chunks.
That creates a grid with 3f+1 rows and 2f+1 columns.
Each row is a primary sliver.
1
10
🧵[15/20]
Secondary encoding, row wise
Start from the original grid, f+1 by 2f+1.
Encode each row from 2f+1 chunks into 3f+1 chunks.
That creates a grid with f+1 rows and 3f+1 columns.
Each column is a secondary sliver.
1
9
🧵[16/20]
So total you have 3f+1 primary slivers, each holds 2f+1 chunks, and 3f+1 secondary slivers, each holds f+1 chunks.
With 3f+1 nodes, node i stores primary i and secondary i.
1
7
🧵[17/20]
Read protocol
To retrieve the file you can use either primary or secondary.
Get f+1 primary slivers and reconstruct, or get 2f+1 secondary slivers and reconstruct.
Overall read cost is comparable to 1D erasure coding. In 1D you split the original data into f+1 big chunks, but in Red Stuff you split the original data into (f+1)(2f+1) smaller chunks. So more chunks does not mean more bytes. The chunks are smaller.
1
8
🧵[18/20]
🧙🪄🧙♂️Recovery protocol, the magic
If a node loses its primary sliver, which has 2f+1 chunks, it asks 2f+1 helper nodes.
Each helper compresses its secondary sliver into one chunk for the target. It sounds like magic, but it is really just linear algebra, so let’s ignore the details here.
After receiving 2f+1 chunks, the node reconstructs its 2f+1 chunk primary sliver.
So it downloads about the same amount of data as it lost. Finally.
Same idea in the other direction.
Lose a secondary sliver, ask f+1 helpers, each compresses a primary sliver into one chunk, then recover the secondary sliver.
1
6
🧵[19/ 20]
So overall.
Secondary slivers heal primary slivers.
Primary slivers heal secondary slivers.
Storage overhead intuition.
Walrus stores two encoded datasets, primary plus secondary, so total replication is about 4.5x for large f.
So compared to the simple 1D scheme with about 3x overhead, Walrus uses about 4.5x storage in exchange for much cheaper healing.
This matters in real networks where nodes churn and crash all the time.
1
7
🧵[20/ 20]
Pheww, finally we have covered a brief intro to Walrus and what makes it different from current decentralized storage designs.
If you want a more detailed explanation, please check the Overleaf version.
overleaf.com/read/tzqrbbwqwp…
Thanks for reading!
If you found this thread useful, please like, retweet, and follow so more people can learn what is behind blockchain.
Comments or questions? Drop them below. ❤️
If there is any protocol or topic you want me to research, let me know 🧑🔬🧪
Keep learning so you will not fork it up 🍴
See you in the next thread.
1
27
Don't know why the link doesn't work
it should be the overleaf web then follows by /read/tzqrbbwqwpyk#faeb58
Feb 6, 2026 · 9:42 AM UTC
19



