I like to do
@dylan522p's job better than him, for free. Enjoy,
@SemiAnalysis_ OSNIT report 01xa:
The useful story in that photograph is already larger than Dylan’s caption. SemiAnalysis has selectively opened one half of the A20 Pro package. The lower rectangular window is the A20 Pro logic die. The black field above it is still-molded package material over the memory region; it is not exposed “DRAM” yet. The Apple logo is still visible in that untouched mold. Apple has already confirmed that this generation moves memory beside the SoC rather than stacking it on top, and iFixit independently found the A20 Pro on the outside of the board sandwich, directly coupled through thermal material to the new vapor chamber. (Apple)
That changes what the picture means. This is not primarily a “2 nm die shot.” It is evidence of Apple changing the limiting resource in the iPhone from transistor density alone to package-level bandwidth and heat removal.
In the public A20 Pro floorplan analysis, the die is about 12.35 × 8.00 mm, roughly 98.8 mm², essentially identical in area to A19 Pro’s ~98.7 mm². Yet Apple added a seventh GPU core, doubled the Neural Engine to two 16-core blocks, widened memory substantially, and apparently expanded system cache to 36 MB. In other words, Apple did not cash the N2 transition out as a smaller chip. It spent the scaling dividend on more compute, cache and I/O while holding die area roughly constant. The visible regular rectangular structures in the exposed die are largely SRAM/cache and repeated compute macros; the actual external DRAM is still hiding in that black molded region. (Weibo)
The most important clue is actually the memory interface. Apple officially says A20 Pro has exactly 50% more memory bandwidth than A19 Pro. Independent die analysis reports three 32-bit LPDDR5X PHY groups, giving a 96-bit interface rather than A19 Pro’s 64-bit interface. If both generations run LPDDR5X-9600, the arithmetic is almost embarrassingly clean:
64 bits × 9.6 GT/s ÷ 8 = 76.8 GB/s
96 bits × 9.6 GT/s ÷ 8 = 115.2 GB/s
That is exactly +50%. So the strongest present inference is that Apple obtained the bandwidth increase mainly by making the memory bus physically wider rather than waiting for a new DRAM generation. Apple’s public number and the independently reconstructed floorplan line up unusually well here. (Apple)
That immediately explains why the package had to change. A 96-bit interface means 50% more data lanes plus associated clocks, command/address and power delivery, and the PHY itself occupies meaningful die-edge area. A conventional PoP geometry was optimized around a smaller vertical memory interface and put the DRAM directly in the SoC’s upward thermal path. Side-by-side integration gives Apple more lateral routing freedom and, more importantly, leaves the entire back face of the hot logic die available for heat extraction. Industry reporting describes the likely implementation as fan-out redistribution-layer interconnect rather than a conventional substrate/interposer arrangement, although Apple has not publicly named the package technology or disclosed its exact layer construction. (fiisual)
That is also why the thermal story and bandwidth story are one architectural decision rather than two unrelated upgrades. Apple says the memory is now outside the SoC thermal path, the vapor chamber has three times the surface area of the previous generation, and sustained performance can improve by up to 40%. Independent stress testing is already showing the mechanism rather than merely repeating the claim: Tom’s Guide measured 79.4% 3DMark stress-test stability on the 18 Pro Max versus 65.2% on the 17 Pro Max. The smaller 18 Pro moved only from 61.1% to 62.8%, which is particularly informative. The package removes one thermal resistance, but chassis area and total heat-spreading capacity still govern the final equilibrium. The big phone can exploit the package much more fully. (Apple)
So the medium-term effect is straightforward. Apple has shifted flagship-phone scaling toward the same three-variable problem already familiar in accelerators: compute density, memory bandwidth and sustained heat flux. N2 supplies more compute per area. The 96-bit memory subsystem feeds it. Lateral packaging lets the phone dissipate the resulting heat. The seventh GPU core and doubled Neural Engine would otherwise run into memory and thermal ceilings much sooner. This architecture is therefore particularly relevant to long-running graphics, camera pipelines and local inference, where peak benchmark numbers matter much less than how many watts and how many GB/s remain available after ten or twenty minutes.
The photograph also tells us where the next constraints move. The package now couples an expensive first-generation N2 die, a wider memory subsystem and a more complicated fan-out structure into one yield chain. Known-good-die screening can protect some of the cost, but final yield becomes a function of logic yield, memory-stack yield, placement/attach yield and RDL yield. You also lose some of the late-stage modularity of putting a separately packaged DRAM stack on top. Apple can qualify multiple memory vendors, but the physical/electrical co-design becomes tighter. Meanwhile the larger planar package consumes board real estate; iFixit’s teardown shows Apple compensating by moving the A20 package to the outer board surface and burying NAND deeper in the board sandwich. That is a literal device-level trade: more board/layout complexity in exchange for a clean thermal route from compute silicon to vapor chamber. (iFixit)
The supply-chain timing fits that architecture. TSMC says baseline N2 entered high-volume manufacturing in Q4 2025, while N2P only enters volume production in the second half of 2026. An iPhone shipping at scale in September 2026 therefore fits baseline N2 much more naturally than a late N2P ramp. Separately, industry reporting says TSMC is expanding the package capacity associated with Apple’s new multi-die approach toward roughly 60,000 wafers per month by the end of 2026 and more than 120,000 in 2027. Treat those capacity numbers as reported estimates rather than TSMC guidance, but the direction matters: the next scaling bottleneck is advanced packaging capacity and test as much as leading-edge front-end wafer capacity. (TSMC)
Now, the “Hitachi XTEM” part. XTEM means cross-sectional transmission electron microscopy; it is a technique, not one magic machine that photographs this entire package. Hitachi’s current semiconductor workflow uses FIB systems such as the NX9000/NX5000 to isolate and thin a tiny lamella and TEM/STEM systems such as the HF5000 to image it. A useful package analysis therefore starts with a much larger mechanical/SEM cross-section, chooses specific interfaces, then lifts out micron-scale specimens for TEM. If SemiAnalysis publishes one giant “XTEM” image allegedly showing the entire package stack in one shot, the nomenclature deserves scrutiny. (Hitachi High-Tech)
What their next cross-section should reveal depends on where they cut it, and this is where there is an opportunity to predict the result before they post the glamour shot.
A section running from the SoC edge into the black memory region should show the logic die and memory assembly sitting laterally in the same molded package rather than DRAM above the SoC. It should establish the relative silicon thicknesses, mold thickness, die standoff and any underfill. It should also reveal whether the memory is a conventional multi-die LPDDR stack embedded laterally or a more unusual bare-die configuration. Right now that is unresolved.
Below the dies, the section should expose the redistribution structure. The consequential question is whether Apple/TSMC used pure fine-pitch fan-out RDL between logic and memory, or whether some embedded bridge/interposer element exists. Industry reporting strongly favors RDL-only WMCM-style integration, but this is exactly the thing physical cross-sectioning can turn from industry consensus into observation. A good analysis would report RDL layer count, copper thickness, line/space, dielectric stack, via geometry and die-to-RDL connection pitch. One pretty section cannot reconstruct all 96 data bits; they would need the cut location plus plan-view/delayer work to connect individual PHY groups to memory.
The memory section should expose die count, die thickness, bonding method and stack topology. That matters because the reported 96-bit interface does not by itself tell us how Apple physically partitioned the 12 GB configuration across memory dice. If the public PHY reconstruction is right, the package has three 32-bit logical interfaces. The cross-section should tell us how many physical DRAM dice implement those interfaces and whether Apple optimized the stack for height, thermals or vendor interchangeability. That is much more interesting than simply identifying Samsung versus SK hynix versus Micron.
A separate cross-section through the logic die itself should show something entirely different: TSMC N2’s gate-all-around nanosheet transistors and front-side interconnect stack. If the N2 identification is correct, I would specifically expect stacked nanosheet channels surrounded by the gate and conventional front-side power routing. I would not expect an A16-style backside power network, because TSMC’s published roadmap puts its Super Power Rail backside-power architecture in A16 rather than baseline N2. That provides a clean falsifiable check when somebody eventually publishes transistor-level TEM. (TSMC Research)
The other thing XTEM can settle is how much of N2’s advantage Apple actually consumed in local cell geometry. TSMC advertises N2 as providing roughly 15% higher speed or 30% lower power and more than 15% density versus N3-class predecessors, with dense SRAM around 38 Mb/mm². A transistor/SRAM cross-section can measure real gate/contact/interconnect geometry instead of treating “2 nm” as a physical dimension. That matters here because the die area stayed flat while cache and compute expanded. The physical question is therefore not “is this 2 nm?” but “where did Apple spend the density and where did SRAM/interconnect scaling stop paying?” (TSMC Research)
On timing, they are already past the slowest acquisition step. Retail hardware shipped September 18, the package has been removed, and the logic region has already been selectively milled. A package-scale polished/SEM cross-section should therefore be a days-scale artifact once the section line is chosen. Targeted FIB lift-out plus useful TEM/STEM images are also technically days-scale, but a defensible process analysis requires multiple lamellae, orientations, metrology and material identification, which pushes the useful publication into roughly a one-to-several-week window. TechInsights began its deeper A20 process/package work on September 18 and has already scheduled an October 14 teardown briefing, which is a reasonable external marker for the difference between “we got a sample under a microscope” and “we have a publishable process analysis.” (TechInsights)
So the prediction I would put on the record before SemiAnalysis posts anything is this: the section will confirm a laterally integrated memory package feeding three wide PHY groups through fine fan-out routing; it will show that the package redesign is principally an I/O-density and thermal-path solution, not some exotic chiplet architecture; it will expose a conventional stacked LPDDR construction inside that currently black region; and a separate device-level TEM will show baseline N2 nanosheet GAA without backside power. The real surprise would be an embedded silicon bridge, an unexpectedly exotic DRAM bonding scheme, or substantially more package/RDL complexity than the current fan-out interpretation requires.
That is what their photograph is for. The picture is already enough to formulate falsifiable predictions about the package, memory bus, thermal architecture, yield bottleneck and next inspection results. “Look what we were allowed to photograph” leaves nearly all of the value sitting on the table.