Senior Engineer @Qualcomm What makes a system well-engineered?

void*
Filter
Exclude
Time range
-
Minimum likes
Replying to @NTDEV_
Can't wait 👀
1
194
The best article!
Hello you fine Internet folks, Today we are looking at Qualcomm's System Level Architecture found in the Snapdragon X2 Elite Extreme SoC and how it handles memory access from both the CPU and GPU side. Hope y'all enjoy! chipsandcheese.com/p/qualcom…
5
909
Replying to @_alwaysalexandr
Interesting. But, Isn't it the same thing but with additional checks and capabilities?
48
Replying to @_alwaysalexandr
"Do even threads require in os dev." If not threads then what? Also, you're talking about Kernel to usermode boundary. Not about threads
1
62
Replying to @VazeKshitij
FSMs are used everywhere. For example in OS, it's used in thread life cycle tracking, when threads move from ready->running or any other state.
2
3
197
Is AI good at multi threading? Personal experience is that I had to clean up a lot of mess it created. It's not bad but it increases the chances of race conditions by a lot.
2
6
693
Replying to @RLeRoux64
Thanks!
1
35
Touched some grass this weekend :)
3
25
1,099
I always wondered what Priority, Preemption, QoS, CPU Affinity and NUMA actually mean when it comes to scheduling. We hear these terms separately, but they all somehow meet when the OS has to answer a simple question: “this thread is ready, where and when should I run it?” I’m going down this rabbit hole now :)
35
1,259
Replying to @marklucovsky
Thank you! I still need to learn more about scheduling with affinity, QoS, NUMA etc.
1
2
122
Replying to @marklucovsky
Is it still the same case with real time priority?
2
191
It’s crazy how advanced modern OS schedulers are. When a thread is ready, the scheduler doesn’t just find a free CPU and run it. It has to think about priority, CPU load, affinity, which core the thread ran on before, cache locality, CPU topology, and on heterogeneous systems, which type of core is a better fit. Move the thread too much and you can lose cache locality. Keep it on one busy core and you lose performance. Put background work on a high-performance core and you may waste power. And these decisions happen continuously while hundreds or thousands of threads are waking, sleeping, blocking and competing for CPU time. Scheduling today is as much about power and hardware topology as it is about sharing CPU time.
4
11
277
12,387
Didn't know that. Just read articles on it. Thanks!
1
1
165
One surprising fact about the first Windows NT: NT meant “New Technology,” and NT 3.1 wasn't built just for Windows apps. It was designed with multiple environment subsystems, supporting Win32, POSIX and character-based OS/2 apps on top of the same NT architecture. Windows was basically one personality on top of NT. Pretty wild for 1993.
Say what you will about Windows (OS), but the NT kernel really is an engineering marvel that still puts Linux to shame in many ways. The quickest way to describe it for a programmer, is NT was more like an object-oriented language, with a strong security model from day one, whereas Linux is very…not. A lot of the “good” features in Linux feel bolted-on (SELinux, Capabilities, Namespaces) because…well they were. I love to imagine an alternate history where NT won. IMO, Microsoft *should* have made an “Open NT” in the early 2000s; not fully GPL-style open, but one where a large org could say…swap out a memory allocator for their own. (they sorta did this with limited source access, but it was too restrictive) You could imagine say…an early Amazon forking OpenNT to create an “AmazonNT” for EC2, where they have a modified scheduler, network stack, whatever. But, the security+compatibility contract keeps a stable baseline on the Microsoft side. Controversial take, but if we enter this era where users are giving AI agents increasingly higher levels of access control; the Linux kernel is legitimately a poor fit. Think about it; answer the question “What exactly is this AI agent allowed to do?” On Standard Linux, it’s disgustingly messy with lots of overlap. Do you use UIDs? GIDs? ACLs? CGROUPs? Policies? SELinux? Filesystem modes? There’s not a singular coherent graph of capabilities you can point to. Too many ways you can escape an initially narrow scope. NT, by comparison, can go the route of explicitly typed resources, and then you could have these really strong centralized audit trails when an agent goes haywire…etc. I know I’m rambling, but the point is…if you were greenfielding an OS kernel from scratch, in 2026, with the intent of being forward-looking, it would *not* look like Linux. Frankly, it’d probably look a lot closer to NT, or even a BSD fork…
10
8
145
47,978
Replying to @lauriewired
Exact thoughts !! What's your take on ReactOS ?
15
1,113
Yes, this is the same account reported multiple accounts.
20
"yashgyy" - it seems like that account is compromised already, that's the reason.
40
You probably keep seeing 40 TOPS, 50 TOPS, 80 TOPS whenever companies talk about NPUs. But what exactly is TOPS? TOPS = Trillion Operations Per Second. Very simply, it is a measure of the peak compute throughput of an NPU under a particular datatype and counting convention. AI workloads involve a huge amount of matrix math, where multiply-accumulate operations (MACs) are extremely common. Think of: "a × b + c" That's one multiply + one addition. When the multiply and addition are counted separately, one MAC is counted as 2 operations. So a common peak calculation is: "TOPS = 2 × MAC units × frequency / 10¹²" That's how thousands of compute units working in parallel can reach trillions of operations every second. So 50 TOPS means a stated peak throughput of 50 trillion operations per second under the precision and counting method used for that specification. But here's the important part: A 50 TOPS NPU isn't automatically faster than a 40 TOPS NPU. You still need to ask: → Is it INT8, INT4, FP16, etc.? → Is the number dense TOPS or sparse TOPS? → Can the workload actually keep the NPU busy? → Is memory bandwidth/data movement becoming the bottleneck? → How efficiently can the compiler/runtime map the model onto the hardware? Precision matters a lot too. Lower-precision operations require less data and can often allow the hardware to execute more operations per second. So 50 INT8 TOPS and 50 INT4 TOPS are not automatically equivalent measurements of capability. And even two NPUs with the same INT8 TOPS can perform very differently on the same model. Because TOPS tells you about peak compute throughput, not guaranteed application performance. Real AI performance also depends on model architecture, memory bandwidth, data movement, software/runtime efficiency, power limits and thermals. So whenever you see: NPU: 50 TOPS read it as: “This NPU has a stated peak throughput of 50 trillion operations per second under a particular precision and counting convention.” Not: “This NPU will always be faster.”
1
24
1,257
KUBSAN → catches C/C++ undefined behavior: invalid shifts, signed overflow, alignment violations, certain invalid casts/operations, etc. This is a good sign.
Kernel UndefinedBehaviorSanitizer (KUBSAN) is now supported on Windows (KubsanInitSystem, etc.)
8
62
3,180
Thanks!
1
83