I love philosophy, all in for finding meaning in and of life with data-if there is any, PhD @vt_ece, postdoc @argonne Views are my own.

Start The MindLab Today
24 citations in the first year. Not bad at all. Gaussian Process still got it. 1/n
5
1
94
10,573
Imagine a model pretrained on natural images being adapted to a medical imaging task. Its frozen features may already capture edges, textures and shapes, while some disease-specific patterns remain poorly represented. ReFine adds a second encoder that learns those missing patterns from the target data, then lets the final predictor use both feature sets together. When the pretrained representation is useful, the target encoder only has to learn a smaller residual signal, which can make the learning problem easier. 2/n
1
16
The paper proves that, under its nonparametric assumptions, this combined representation achieves a convergence rate as good as training from scratch up to logarithmic factors, and can approach a near-parametric rate when the source features are highly informative. The authors also test the idea across image, text, tabular and single-cell settings, including adding spatial information to a pretrained single-cell model at adaptation time. The guarantee is specifically about convergence rates under the paper’s theoretical assumptions, while finite-sample behavior still depends on the data and model. The broader idea is compelling: transfer becomes safer when the target task has a dedicated channel for whatever the source representation failed to capture. 3/n
3
"There is a mystical fool in me that proved to be stronger than all my science." Carl Jung
4
204
Pooja Algikar retweeted
we're going to need a lot more mathematicians
9
6
41
3,350
"The only intelligent tactical response to life's horror is to laugh defiantly at it" Søren Kierkegaard
3
52
Pooja Algikar retweeted
my life got much better when I allowed myself to write relentlessly (in small fonts)
1
1
4
133
AI field allows too many customized definitions.
55
Imagine an image model with distinct clusters for birds, bicycles and chairs. XTransfer summarizes each layer’s activations into a compact feature map and pairs sensor classes with suitable image-class clusters. Walking could be assigned to the bird cluster because the matching follows feature geometry. Small trainable connectors then guide sensor features toward their assigned clusters while keeping different activities apart. The original layers stay frozen during this process, keeping those reference points stable. 2/n
1
27
XTransfer also searches across pretrained networks and assembles useful layers into a compact sensor model. It selects layers that help separate the target classes while staying within the device’s computing and memory budget. In the authors’ single-source deployment tests, the resulting models ran 1.4 to 4.1 times faster than the ResNet18 baseline. The class-pairing procedure requires labeled target examples and enough source classes to provide distinct reference clusters. Together, these ideas offer a practical way to build small sensor models using both the learned geometry and individual layers of existing networks. 3/n
17
life full of surprises. everywhere. these days. this sentence is not from a fiction book: "where refers to the source, neuron, or some tree in a random forest in which transfer learning happens."
2
174
Pooja Algikar retweeted
A day will come when the hard problem of consciousness will really be the hard problem of AI.
1
2
220
The important result is that the authors can characterize exactly when this sharing actually helps. The estimation error effectively separates into the error from learning the shared component, where you benefit from the full dataset, and the error from learning the smaller group-specific correction. If those pieces are structurally simpler than the full target function, the transfer estimator can converge faster than training independently on each group. With deep ReLU networks, they further show that if the underlying functions have a hierarchical compositional structure, complex functions built from simpler low-dimensional pieces, the convergence rate depends more on that intrinsic structure than on the raw input dimension, preserving one of the reasons deep networks can cope with high-dimensional data. 2/n
1
1
1
55
This paper puts a surprisingly clean mathematical story underneath a very common ML recipe: learn what is common from lots of data, then learn only what changes using little data. Many real transfer-learning problems are exactly of this form: multiple related populations with strong shared structure but limited samples for each individual population. It also gives a useful warning: transfer is not automatically beneficial; the theory relies heavily on the assumption that the target functions can be decomposed into a shared component plus a relatively simple group-specific offset. If groups differ in a fundamentally non-additive way, or the shared structure is weak, the guarantees no longer say that transfer should help. 3/n
1
1
24