The paper proves that, under its nonparametric assumptions, this combined representation achieves a convergence rate as good as training from scratch up to logarithmic factors, and can approach a near-parametric rate when the source features are highly informative. The authors also test the idea across image, text, tabular and single-cell settings, including adding spatial information to a pretrained single-cell model at adaptation time. The guarantee is specifically about convergence rates under the paper’s theoretical assumptions, while finite-sample behavior still depends on the data and model. The broader idea is compelling: transfer becomes safer when the target task has a dedicated channel for whatever the source representation failed to capture. 3/n