A high performance, homegrown Solana validator Developing the full node client Mithril Team: @7LayerMagik & @dubbel06

Based in United States
Filter
Exclude
Time range
-
Minimum likes
We would be interested in getting the mechanism audited for everyone's benefit, but we need a sugar daddy for that
5
6
62
The multicore bench stats in the original thread were outdated and have now been updated
Replying to @OverclockSol
9/ Across 8 cores we measured 1.217M signatures/s, short of 8× the single-core rate because all-core AVX-512 drops the clocks.
1
3
25
5,674
Thank you to @Allnodes for giving us a free Ryzen server rental to test these Zen5 benchmarks on!
1/ Ed25519 verification is a major cost in Solana block replay. Mithril is written in Go, and the standard Go library wasn't enough for the throughput we wanted. Today we're releasing narya-ed25519. 5.8× faster than crypto/ed25519 at 4.7 µs/sig (Zen5). github.com/Overclock-Validat…
1
1
24
1,161
11/ Narya is alpha and unaudited, and the AVX-512 backend is opt-in rather than automatic. We differential-test against Go's stdlib and curve25519-voi, run the CCTV and Wycheproof suites, and fuzz, but it still needs independent review. Apache-2.0, attributions in NOTICE.
1
7
387
10/ Regarding our work here, we want to say thank you to our AI overlords. AI was deeply involved in the R&D (Codex, ChatGPT Pro, Claude) and helped explore code, analyze profiles/math, generate hypotheses and tests, write docs, and implemented all the code including assembly.
1
7
421
9/ Across 8 cores we measured 1.217M signatures/s, short of 8× the single-core rate because all-core AVX-512 drops the clocks.
1
7
1,499
8/ For recurring keys it goes further, caching the decoded public key and then promoting hot ones to a precomputed table. With 64 promoted keys on the same Zen 5 host, n=64 measured 3.776 µs/sig, 7.3× stdlib. How much you get depends on key population and locality.
1
6
388
7/ You need several signatures in hand at once, which replay gives you. They don't have to be related, though: cold verification needs no cached or precomputed public-key state, so any unrelated signatures can fill the lanes.
1
6
394
6/ That independence is deliberate. The faster option is aggregate batch verification, folding every signature into one randomized equation, but it returns a single verdict for the whole set. A validator has to reject a transaction, not a block.
1
6
431
5/ Narya also vectorizes across signatures rather than within one: a whole signature per lane, eight per ZMM register, which is why everything in the code is named x8. Every lane has its own equation, so you get eight separate verdicts, identical to verifying them individually.
1
7
481
4/ The backend we ship uses five 51-bit limbs instead. That leaves headroom for unreduced sums to feed straight into IFMA source operands, which the point formulas exploit constantly, saving a carry pass on nearly every one.
1
7
521
3/ We started from Firedancer's 2023 AVX-512 work, which represents a field element as six limbs of radix 2^43. We ported it to Go first and kept it as a correctness reference for differential testing. github.com/firedancer-io/fir…
1
1
10
727
2/ Narya is a no-cgo Go library using AVX-512 IFMA assembly. The numbers below are cold strict verification on 1232-byte messages, with no per-key state cached between calls. Most of the gain shows up once you have eight signatures, since that's what fills the lanes.
1
8
936
1/ Ed25519 verification is a major cost in Solana block replay. Mithril is written in Go, and the standard Go library wasn't enough for the throughput we wanted. Today we're releasing narya-ed25519. 5.8× faster than crypto/ed25519 at 4.7 µs/sig (Zen5). github.com/Overclock-Validat…
7
22
100
121,277
Replying to @toly
Okay, Mithril validator is running at less than 3GB RAM on the Alpenglow cluster with a small Ryzen. We just need some of your state compression ideas implemented so that we don't need to plug a NVME into the phone when we have Mainnet state to deal with. The massive 50k tps synthetic blocks on ag cluster will also be challenge for a phone...
1
2
68
Replying to @toly
Is this in the plans at all? @bw_solana
1
3
71
Replying to @BloodReaver
The real hardware requirement comparison will be once our Alpenglow version hits Mainnet... but it should still be good. Agave's hardware requirements have also trended down over time by quite a bit. On Alpenglow testnet cluster, Agave is using around 6 GB RAM last time we saw.
5
36
Replying to @BloodReaver
For something that looks more like Mainnet traffic: 8 GB RAM, 6 fast cores (newer Ryzen cpu's), and 1 TB NVME. For Alpenglow testnet, you don't need nearly as much storage space (AccountsDB and snapshots are very small), but they blast it at times with over 50k tps synthetic traffic and a machine with just a few cores will struggle greatly. We need to improve our sigverify implementation quite a bit though so we will see how things go after that.
2
4
178
6/ And with that, we'd be happy to see more contributors join the fellowship ⚒️ If you're interested, join our Discord and send word in our Mithril channels: discord.gg/sHzb3EvmkR
6
294
5/ Core validator features left to implement include: snapshot production, Turbine retransmit (ingestion is working), other networking and repair updates, broader validator hardening, and block builder integrations, among other things.
1
3
301