Sharing insights on Probability, Statistics, ML, DL and AI research. Subscribe for recent research paper discussions at $2/month. DM to collaborate.

Optimal Transport for Network Comparison: Optimal Transport (OT) provides a principled way to compare networks by treating them as distributions of nodes, edges, features, or embeddings rather than comparing them only through graph-level statistics. Given two probability distributions μ and ν, OT seeks the cheapest way to transform μ into ν: W_c(μ,ν) = inf₍π∈Π(μ,ν)₎ 𝔼₍(x,y)∼π₎[c(x,y)], where c(x,y) is the cost of transporting mass from x to y and Π(μ,ν) is the set of couplings with marginals μ and ν. For networks, nodes can be represented by feature vectors, degrees, centralities, or learned embeddings. OT then finds a soft correspondence between nodes of two networks, making it useful even when the networks have different sizes or no obvious node-to-node alignment. A major advantage is that OT compares the geometry and distribution of network structure, rather than relying solely on handcrafted summary statistics. Entropic regularization, Wε(μ,ν) = minπ∈Π(μ,ν) ⟨π,C⟩ + ε∑ᵢⱼ πᵢⱼ(log πᵢⱼ−1), also makes computation substantially faster via Sinkhorn-type algorithms. In Statistics, OT enables distributional comparison, clustering and hypothesis testing for network populations. In ML, it supports graph matching, domain adaptation, graph representation learning and generative modeling. In AI, it can compare knowledge graphs, social networks, molecular graphs and neural representations. The key idea is simple: instead of asking whether two networks have the same nodes, ask how much “work” is required to transform the structure of one network into the other.
2
33
239
15,787
Congratulations @deepigoyal on @temple :) More researchers should be celebrated like this!
India has always idolised actors and cricketers. The last few years, entrepreneurs and founders have also gotten mobbed at the airports. How I wish we started celebrating scientists as much as we celebrate actors, cricketers, and founders. On that note, meet the Temple scientists behind Brain Flow: Dr. Divya Gulati (IISc), Dr. Sanchit Gupta (IISc), Dr. Yukti Chopra (SISSA Trieste), and Nitish Kumar (IIT Bombay). This team led, built, and validated the first continuous, wearable measurement of blood flow to the brain. Your brain uses 20% of your blood supply. It controls your hormones, sleep, immunity, and repair. That supply drops every year as you age, and you have never measured it. Even hospitals can only capture a single snapshot. Measuring and maintaining brain flow could be one of the most important health and longevity interventions of our time. At Temple, we have some of the best minds in the country, working on the one problem that sits upstream of every other: extending healthy human life. I am proud of our team, and looking forward to a lot of firsts and breakthroughs in the future.
12
3,844
ML MCQ: Suppose a model minimizes L(θ) = MSE(θ) + λ‖θ‖₂² If λ is increased substantially, what is the most likely effect? A) Training error decreases and variance increases B) Training error increases and model variance decreases C) Both training error and variance increase D) The model becomes more complex
5
2
67
12,975
LASSO (Least Absolute Shrinkage and Selection Operator) is a regularization algorithm for linear regression that performs prediction and variable selection simultaneously. Given observations (xᵢ,yᵢ), the LASSO estimator solves β̂ = argminβ { (1/2n)∑ᵢ(yᵢ − xᵢᵀβ)² + λ‖β‖₁ }, where λ ≥ 0 controls the strength of regularization and ‖β‖₁ = ∑ⱼ|βⱼ| is the L₁ norm. Unlike ordinary least squares, the L₁ penalty can force some estimated coefficients exactly to zero, producing a sparse model. The geometric shape of the L₁ constraint explains this behavior: its corners make solutions more likely to lie on coordinate axes. LASSO therefore provides a principled approach to high-dimensional prediction when many candidate variables may be irrelevant. It is particularly important when p, the number of predictors, is comparable to or substantially larger than n, the sample size. LASSO has applications in statistical machine learning, including feature selection, high-dimensional regression, genomics, signal processing and compressed sensing. Its extensions—such as Elastic Net, Group LASSO and adaptive LASSO—allow different forms of structural sparsity. The method beautifully connects statistical regularization, convex optimization and the modern theory of high-dimensional inference.
1
19
176
9,355
DBSCAN (Density-Based Spatial Clustering of Applications with Noise) is an unsupervised learning algorithm that identifies clusters as regions of high data density rather than assuming a particular cluster shape. Given a radius ε and a minimum number of points MinPts, an observation x is considered a core point if its ε-neighborhood contains at least MinPts observations. Clusters are then formed by connecting density-reachable points, while isolated observations can be classified as noise. Unlike K-Means, DBSCAN does not require the number of clusters to be specified in advance and can discover clusters with highly non-convex geometries. Its behavior is governed by the local density structure of the data, making it closely related to ideas from nonparametric statistics and density estimation. DBSCAN has applications in statistical machine learning including spatial data analysis, anomaly detection, image segmentation, customer behavior analysis and geospatial pattern recognition. From a statistical perspective, its notion of local density provides a simple way to distinguish structured populations from outliers. It is particularly useful when clusters have irregular shapes or when observations that do not belong to any meaningful population should be explicitly identified as noise. More recent density-based methods, such as HDBSCAN, extend these ideas to datasets with varying density levels.
26
184
7,126
ML MCQ: In a decision tree, increasing the maximum depth of the tree generally causes: A) Higher training error and higher bias B) Lower training error and potentially higher variance C) Lower variance and higher bias D) No change in the model's complexity
5
2
23
5,813
Metropolis–Hastings (MH) is a Markov chain Monte Carlo (MCMC) algorithm for generating samples from a complicated probability distribution when direct sampling is difficult. Given a target density π(x), the algorithm proposes a new state x′ from a proposal distribution q(x′|x) and accepts it with probability α(x,x′) = min{1, [π(x′)q(x|x′)]/[π(x)q(x′|x)]}. If the proposal is rejected, the chain remains at x. Under suitable conditions, the resulting Markov chain has π as its stationary distribution, so averages over the simulated samples can approximate expectations under the target distribution. MH is fundamental to statistical machine learning, particularly Bayesian inference. When the posterior distribution p(θ|D) is analytically intractable, MCMC can generate approximate posterior samples, allowing us to estimate parameters, credible intervals and predictive distributions. Applications include Bayesian regression, hierarchical models, latent-variable models, graphical models and uncertainty quantification. The algorithm also illustrates a central idea in computational statistics: rather than solving an integral or optimization problem directly, construct a stochastic process whose long-run behavior reproduces the desired distribution. Modern methods such as Hamiltonian Monte Carlo and Langevin Monte Carlo build on this principle.
3
29
188
6,962
Gaussian Processes (GPs) provide a powerful framework for Bayesian nonparametric regression and classification. Instead of assuming a finite-dimensional parameter vector, a GP places a probability distribution directly over functions: f(x) ∼ GP(m(x), k(x,x′)), where m is the mean function and k is a covariance kernel describing how function values at different inputs are related. Given noisy observations yᵢ = f(xᵢ) + εᵢ, the joint Gaussian structure allows the posterior distribution of f at new points to be computed analytically. The predictive distribution has the form f(x*) | X,y,x* ∼ N(μ*, σ*²), providing both a prediction and an explicit measure of uncertainty. GPs have extensive applications in statistical machine learning, including regression, spatial statistics, time-series modeling, Bayesian optimization, surrogate modeling and uncertainty quantification. Different kernels encode different assumptions about smoothness, periodicity and similarity, making kernel selection a form of inductive modeling. Gaussian processes also provide a statistical foundation for understanding uncertainty in machine-learning predictions. Their connections to reproducing-kernel Hilbert spaces, kernel methods and Bayesian inference make them an important bridge between classical statistical modeling and modern machine learning.
11
96
716
20,684
t-SNE (t-distributed Stochastic Neighbor Embedding) is a nonlinear dimensionality-reduction algorithm designed to visualize high-dimensional data in two or three dimensions. Rather than preserving global Euclidean distances, t-SNE attempts to preserve local neighborhood structure. For high-dimensional observations xᵢ, it defines conditional similarities using Gaussian kernels: pⱼ|ᵢ ∝ exp(−‖xᵢ−xⱼ‖²/2σᵢ²). In the low-dimensional representation y₁,…,yₙ, similarities are modeled using a Student-t distribution: qᵢⱼ ∝ (1+‖yᵢ−yⱼ‖²)⁻¹. The embedding is obtained by minimizing the Kullback–Leibler divergence KL(P‖Q) = ∑ᵢⱼ pᵢⱼ log(pᵢⱼ/qᵢⱼ), typically using gradient-based optimization. The heavy-tailed Student-t distribution helps separate points that are moderately distant in the original space while retaining nearby observations together. t-SNE is widely used in statistical machine learning for exploratory data analysis and visualization of high-dimensional datasets, including images, text embeddings, gene-expression data and learned representations from neural networks. It can reveal clusters and nonlinear structure that are difficult to detect using PCA. However, the resulting geometry can depend strongly on hyperparameters and initialization, so apparent distances between clusters should not automatically be interpreted as statistically meaningful global distances.
2
30
249
14,620
AdaBoost (Adaptive Boosting) is an ensemble learning algorithm that combines many weak classifiers to construct a strong predictor. Given training observations (xᵢ,yᵢ), yᵢ ∈ {−1,+1}, AdaBoost assigns a weight wᵢ to each observation. Initially, all observations receive equal weight. At iteration t, a weak learner hₜ is trained using the current weights, and its weighted error is εₜ = ∑ᵢ wᵢ 𝟙[hₜ(xᵢ) ≠ yᵢ]. The learner receives weight αₜ = ½ log((1−εₜ)/εₜ), and the observation weights are increased for incorrectly classified observations and decreased for correctly classified ones. The final classifier is H(x) = sign(∑ₜ αₜhₜ(x)). The remarkable feature of AdaBoost is that it transforms a sequence of relatively simple, weak learners into a powerful ensemble by concentrating subsequent learning on difficult observations. From a statistical perspective, AdaBoost can be interpreted as a form of stagewise additive modeling that approximately minimizes exponential loss. It connects empirical risk minimization, convex optimization and ensemble learning. Applications include classification, face detection, text categorization, medical diagnosis and pattern recognition. Its underlying boosting principle also inspired modern algorithms such as gradient boosting and XGBoost.
29
190
5,388
K-Means is one of the most widely used unsupervised learning algorithms for partitioning observations into K clusters. Given data x₁,…,xₙ ∈ ℝᵖ, it seeks cluster centers μ₁,…,μₖ that minimize the within-cluster sum of squared distances: min_{μ₁,…,μₖ} ∑ᵢ minₖ ‖xᵢ − μₖ‖². The algorithm alternates between two simple steps: assignment, where each observation is allocated to its nearest centroid, and update, where each centroid is replaced by the mean of the observations assigned to it. Each iteration decreases the objective function, although the algorithm can converge to a local rather than global optimum, making initialization important. K-Means has numerous applications in statistical machine learning, including customer segmentation, image compression, document clustering, anomaly detection and exploratory data analysis. Statistically, it can be interpreted as estimating a finite mixture structure under particular assumptions, and it is closely related to vector quantization and the Expectation–Maximization algorithm for Gaussian mixture models. Its simplicity also makes K-Means a useful starting point for understanding more sophisticated clustering methods and the broader principle of optimizing an empirical objective to uncover hidden structure in data.

ALT Image: https://share.google/kWYIGzX40WdRldDg4

1
25
191
6,650
We should not delegate everything to LLMs. We should not stop thinking!
1
1
32
2,738
Principal Component Analysis (PCA) is a fundamental unsupervised learning algorithm for reducing the dimensionality of data while preserving as much variation as possible. Given centered observations x₁,…,xₙ ∈ ℝᵖ with sample covariance matrix S, the first principal component is the direction v₁ solving v₁ = argmax_{‖v‖=1} vᵀSv. Thus, v₁ is the eigenvector of S corresponding to its largest eigenvalue. Subsequent principal components are obtained from orthogonal directions associated with the remaining eigenvalues. Equivalently, PCA can be computed through the singular value decomposition of the data matrix. PCA has extensive applications in statistical machine learning, including dimensionality reduction, visualization, feature extraction, noise reduction and exploratory data analysis. In high-dimensional statistics, PCA is particularly useful for identifying latent low-dimensional structure hidden within noisy observations. It also forms the foundation of methods such as probabilistic PCA and factor analysis. In modern ML, PCA can reduce computational complexity before classification or regression and provide compact representations for downstream models. Its close relationship with covariance estimation, spectral theory and least-squares reconstruction makes PCA a natural bridge between classical multivariate statistics and machine learning.
3
39
187
16,074
Gradient Boosting is an ensemble learning algorithm that constructs a powerful predictor by sequentially adding many weak learners, typically shallow decision trees. Starting with an initial model F₀(x), the algorithm iteratively updates Fₘ(x) = Fₘ₋₁(x) + ηhₘ(x), where hₘ is trained to approximate the negative gradient of the chosen loss function and η > 0 is the learning rate. For squared-error regression, this gradient corresponds to the residuals yᵢ − Fₘ₋₁(xᵢ), so each new tree attempts to correct the errors of the existing model. From a statistical perspective, gradient boosting can be interpreted as functional gradient descent: instead of optimizing a finite-dimensional parameter vector, it performs optimization directly in a space of functions. Regularization through learning-rate shrinkage, tree depth, subsampling and early stopping helps control overfitting. Gradient boosting is widely used in statistical machine learning for regression, classification, ranking and risk prediction. Modern implementations such as XGBoost, LightGBM and CatBoost have become important in structured and tabular-data problems. The method also connects statistical concepts such as empirical risk minimization and additive models with modern computational optimization, making it a particularly useful bridge between classical statistics and machine learning.
1
23
116
4,182
Bayesian Optimization (BO) is a machine-learning framework for optimizing expensive or black-box functions when evaluating the objective is costly. Suppose we want to maximize an unknown function f(x), where each evaluation may require a costly experiment or simulation. BO places a probabilistic surrogate model—often a Gaussian process—on f: f ∼ GP(m(x), k(x,x′)). After observing data Dₜ = {(xᵢ,f(xᵢ))}₁ᵗ, the Gaussian process produces a posterior predictive distribution for every candidate x. Instead of blindly evaluating points, BO constructs an acquisition function that balances exploration of uncertain regions with exploitation of locations predicted to have high values. Popular choices include Expected Improvement, Probability of Improvement and Upper Confidence Bound. The algorithm has important applications across statistical machine learning, including hyperparameter optimization, experimental design, robotics, drug discovery and engineering. In statistics, BO can be viewed as sequential decision-making under uncertainty: each new observation updates the posterior distribution and determines where to sample next. Its mathematical foundations connect Gaussian processes, Bayesian inference, stochastic optimization and information theory, making BO a powerful example of how probabilistic modeling can guide efficient learning when data collection itself is expensive.
8
40
234
9,713
Support Vector Machine (SVM) is a powerful supervised learning algorithm designed to find decision boundaries that maximize the separation between classes. For binary labels yᵢ ∈ {−1,+1}, the soft-margin SVM solves min_{w,b,ξ} ½‖w‖² + C∑ᵢξᵢ subject to yᵢ(wᵀxᵢ+b) ≥ 1−ξᵢ, ξᵢ ≥ 0. The parameter C controls the trade-off between maximizing the margin and penalizing classification errors. Only observations close to the decision boundary—the support vectors—directly determine the fitted classifier. Through the kernel trick, SVMs can implicitly map observations into high-dimensional feature spaces using kernels K(x,x′), allowing nonlinear decision boundaries without explicitly constructing the transformation. Common kernels include polynomial and Gaussian RBF kernels. SVMs have important applications in statistical machine learning, including text classification, image recognition, bioinformatics, handwriting recognition and high-dimensional prediction. From a statistical perspective, the large-margin principle provides a form of regularization that can improve generalization, while the kernel formulation connects SVMs to reproducing-kernel Hilbert spaces and nonparametric estimation. Thus, SVMs provide a beautiful example of how convex optimization and statistical learning theory combine to construct flexible predictive models.
12
117
3,554
Random Forest is an ensemble learning algorithm that combines many decision trees to produce a more stable and accurate predictor. Given training data {(xᵢ,yᵢ)}₁ⁿ, each tree is trained on a bootstrap sample of the observations, while considering a random subset of features at every split. If T₁(x),…,T_B(x) are the resulting trees, the forest prediction for regression is f̂(x) = 1/B ∑ᵦ Tᵦ(x), while classification typically uses majority voting. The key statistical idea is variance reduction: individual decision trees can have high variance, but averaging many sufficiently diverse trees substantially improves generalization. Random forests are widely used in statistical machine learning because they can model complex nonlinear relationships and high-order interactions without requiring a parametric specification. They are applied to classification, regression, feature selection, anomaly detection and risk prediction. Their built-in estimates of variable importance can help identify influential predictors, while out-of-bag observations provide a convenient form of internal validation. Random forests are particularly valuable in high-dimensional problems where linear models may be too restrictive and where interpretability of a single decision tree is traded for the predictive robustness of an ensemble.
2
27
164
6,069
Expectation–Maximization (EM) is a fundamental algorithm for maximum-likelihood estimation when a statistical model contains unobserved or latent variables. Suppose X denotes observed data and Z latent variables, with parameters θ. Rather than directly maximizing the often complicated likelihood p(X|θ), EM alternates between two steps. In the E-step, it computes the conditional distribution of the latent variables using the current parameter estimate: Q(θ|θₜ) = E[log p(X,Z|θ) | X,θₜ]. In the M-step, it updates the parameters by maximizing this expected complete-data log-likelihood: θₜ₊₁ = argmaxθ Q(θ|θₜ). The procedure is guaranteed not to decrease the observed-data likelihood under standard regularity conditions, although it may converge to a local rather than global optimum. EM is central to statistical machine learning. Its most famous application is fitting Gaussian mixture models (GMMs), where the latent variable indicates the cluster generating each observation. It is also used in hidden Markov models, missing-data problems, probabilistic PCA, topic models and latent-class models. More broadly, EM illustrates a powerful principle in Stat-ML: complicated inference can often be transformed into a sequence of simpler conditional estimation problems by introducing appropriate latent structure.
2
42
298
11,030
PageRank is a celebrated spectral algorithm that ranks nodes in a directed network according to their structural importance. Given a graph with transition matrix P, PageRank defines a probability vector π satisfying π = αPᵀπ + (1−α)v, where 0 < α < 1 is the damping factor and v is a personalization distribution. Equivalently, π is the stationary distribution of a random walk that follows network links with probability α and jumps according to v otherwise. The damping term guarantees existence and uniqueness of π even when the original network is disconnected or contains dangling structures. At a deeper level, PageRank is an eigenvector problem: π is the dominant eigenvector of αPᵀ + (1−α)v1ᵀ. Its ideas extend far beyond web search. In statistics, PageRank can quantify influence, centrality and dependence in networks, with applications to citation networks, social networks, financial systems and biological interaction graphs. In machine learning, PageRank-like diffusion provides graph-based semi-supervised learning, label propagation, recommendation, anomaly detection and information retrieval. Personalized PageRank is particularly useful for constructing task-specific representations: starting from a seed distribution v, probability mass propagates through the graph according to its topology. Modern graph neural networks and graph-based ranking methods can therefore be viewed, in part, as sophisticated descendants of this fundamental random-walk principle.
3
32
245
11,518
Probability and Statistics retweeted
Luis Martínez-Zoroa and His Contributions to PDE: Luis Martínez-Zoroa is a young mathematician working at the intersection of PDE, harmonic analysis and incompressible fluid dynamics, with a particular focus on singularity formation and loss of regularity. His work has addressed some of the most delicate questions surrounding Euler, Navier–Stokes and Surface Quasi-Geostrophic (SQG) equations. A central theme of his research is understanding how initially smooth solutions can lose regularity. With Diego Córdoba, he constructed global unique solutions of dissipative SQG that undergo instantaneous loss of regularity in the supercritical regime. They also established strong ill-posedness results for SQG in critical function spaces. His recent work has moved toward finite-time singularity formation in three-dimensional fluid equations. With Córdoba and Fan Zheng, he proved finite-time singularities for the 3D incompressible Euler equations through a novel blow-up mechanism. Even more recently, the same collaboration established finite-time blow-up for forced hypodissipative Navier–Stokes equations, showing that classical finite-energy solutions can become singular under fractional dissipation below a critical threshold. Together, these results form an important programme for understanding the fundamental question: Can nonlinear fluid equations evolve from smooth data to singular states in finite time? Martínez-Zoroa's work is particularly notable for developing rigorous mechanisms for regularity breakdown, ill-posedness and singularity formation—central themes in modern nonlinear PDE.
57
267
11,129