@MIT trained Neuroscientist interested #Neural_Nets, #BioAI, #NeuroAI 🧠⚡🤖 substack.com/@bioai2neuro 🎯 aibio28ai@gmail.com

Texas, United States
Ability of a Structural World Model to Detect Cryptic Pockets from Apo Structure biorxiv.org/content/10.64898… Summary: This study presents a method for identifying cryptic drug-binding pockets directly from a single apo protein structure, without requiring molecular dynamics simulations, conformational sampling, co-folding predictions, or external pocket-detection tools. The approach uses a proprietary structural world model that generates a per-residue latent representation from protein coordinates. This latent state is read out as a cryptic-lining score, suggesting the model implicitly captures conformational flexibility that other methods must explicitly simulate. Predicted pockets are constructed as residue sets rather than fixed geometric spheres. High-scoring residues seed candidate pockets, which are expanded into distinct, non-overlapping predictions through a residue-growth strategy and geometric non-maximum suppression. On CryptoBench (231 test proteins), the model achieves 84.8% top-1 and 95.2% top-5 localization accuracy. Similar performance is observed on a CryptoBank subset, with 84.6% top-1 and 99.0% top-5 accuracy. Residue-level classification is strong (AUC = 0.8465), although exact pocket-boundary recovery remains more challenging under stricter overlap criteria. The method generalizes well to unseen proteins, recovering the known WRN helicase allosteric site at rank 1 across multiple apo structures after all WRN proteins were removed from training. A key advantage is its ability to place the true cryptic site at rank 1, addressing a common limitation of ensemble-based approaches. Compared with single-structure baselines such as P2Rank, DeepPocket, PocketMiner, and fpocket, the model achieves substantially higher top-1 hit rates. It also complements the ensemble-based method OpenDDE, identifying many cryptic sites missed by OpenDDE while retaining all of OpenDDE’s top-5 hits. Overall, the work demonstrates that latent representations learned by a structural world model can effectively detect and rank cryptic pockets from apo structures alone, with performance improving as training data increases. #DrugDiscovery #CrypticPockets #BioAI #AIforScience #ProteinStructure
2
2
5
547
A foundation model of vision, audition, and language for in-silico neuroscience arxiv.org/abs/2605.04326 #Foundation_models, #NeuroAI, #NeuralNets
9
43
2,519
Sentience and the Origins of Consciousness: From Cartesian Duality to Markovian Monism mdpi.com/1099-4300/22/5/516
1
5
13
680
Inspired by @SGRodriques’s 12 Millennium Problems in Biology, I put together 20 open problems in neuroscience that I believe are worth tackling over the coming decade. These are meant to spark curiosity, discussion, and new ideas—not to be definitive. I’m a researcher/neurophysiologist sharing this purely as a hobby and for fun. I have no commercial interests and currently represent no company or organization as of now! #Neuroscience #NeuroAI #BioAI #NeuralNetworks #BrainResearch #ComputationalNeuroscience #SystemsNeuroscience #Neurotechnology #Connectomics #Neurobiology #FutureOfNeuroscience #OpenProblems
13
9
34
2,021
The recipe for intelligence in natural and artificial systems cell.com/neuron/fulltext/S08…
5
7
56
3,573
Important read for BioAI community Why machines don’t speak biology: Toward native biological language models cell.com/cell/fulltext/S0092… From AlphaFold to Virtual Cells The recent wave of BioAI has been fueled by the success of foundation models such as AlphaFold, which demonstrated that AI can uncover meaningful biological rules from massive datasets. This success has inspired efforts to build increasingly ambitious models capable of understanding genomes, cells, tissues, and ultimately entire biological systems. The long-term vision is to create "virtual cells" and predictive models that can simulate development, aging, disease, and therapeutic responses. The Challenge of Biological Complexity However, the authors argue that biology is fundamentally different from many problems where AI has excelled. AlphaFold benefited from strong evolutionary constraints and extensive structural data, but moving from proteins to cells and tissues introduces vastly greater complexity. Living systems operate across multiple spatial and temporal scales, with billions of interacting molecules and cells generating behaviors that cannot always be predicted from their individual components. Why Emergence Matters A central theme of the article is that biological systems exhibit emergence—properties that arise from interactions across scales and cannot be fully explained through reductionism alone. While AI can be extremely powerful for predicting patterns and responses within well-defined domains, it may struggle to reconstruct the historical, evolutionary, and stochastic processes that shape biological complexity. Learning Biology Through Processes To overcome these limitations, the authors propose focusing on canonical biological processes such as DNA replication, transcription, the cell cycle, and embryonic development. These reproducible processes provide common spatial and temporal frameworks that enable causal reasoning, integration of diverse datasets, and mechanistic understanding rather than simple pattern recognition. Toward Biological World Models Looking ahead, the authors envision the development of biological world models that combine AI with mechanistic and systems-level biology. Rather than building a single grand unified model of life, future BioAI systems may integrate standardized and causal models across multiple scales—from molecules to cells to tissues—to better predict, explain, and ultimately engineer biological systems.
7
14
112
5,377
Biology may become the most important frontier for AI. The latest @CellCellPress special focus on #BioAI captures a profound shift in science: the transition from observing living systems to computationally understanding, predicting, and eventually designing them. From proteins and cells to brains, aging, and therapeutics, biology is becoming an information science. 🧬🤖 Over the coming days, I'll explore all these papers (in detail) from this collection and unpack the big ideas shaping the future of AI for biology. 🧬🤖🧵 cell.com/cell/current#fullCo… #AIforScience #ComputationalBiology #DrugDiscovery #Neuroscience
3
4
20
1,389
1
166
sex-specific and dimorphic neural networks
1
4
249
Here is the summary of the complete 🧠🪰 fly connectome paper: Sexual dimorphism in the complete Drosophila male central nervous system connectome cell.com/cell/fulltext/S0092… 🧠🪰 The complete male Drosophila CNS connectome maps 166,700 neurons, spanning the brain and nerve cord, with 11,710 neuron types annotated at synaptic resolution. 🔬 For the first time, male and female brain connectomes are compared at synaptic resolution, revealing: • 8,069 isomorphic neuron types • 138 dimorphic types • 289 male-specific types • 71 female-specific types 1. One striking finding: sex-specific and dimorphic neurons are concentrated in higher-order brain centers, while sensory and motor peripheries remain largely conserved. 2. The connectome also enables complete sensory-to-motor circuit analysis, spanning visual, auditory, olfactory, and taste pathways. 3. Most importantly, dimorphic neurons can reroute information differently between sexes, providing a circuit-level framework for understanding how genetic sex differences shape behavior. This is more than a wiring diagram—it is a foundation for building computational models of behavior from neurons → synapses → circuits → action. 🧠⚡🤖 #Neuroscience #Connectome #NeuroAI #Drosophila #BrainScience #AI #ComputationalNeuroscience
3
3
26
2,220
Brain Foundation Models: A Survey on Advancements in Neural Signal Processing and Brain Discovery arxiv.org/html/2503.00580v1
2
11
41
2,814
This argument remained a talk my one of my scientific heroes, while I was a postdoc at @MIT Please check this Cornelia (Cori) Bargmann's talk: piped.video/watch?v=138NzNhs…
2
45
Youth-associated protein TIMP2 regulates microglial state and function in healthy and aged mice nature.com/articles/s41467-0…
2
1
15
940
Learning functional properties of proteins with language models nature.com/articles/s42256-0… Protein property prediction has increasingly relied on data-driven machine learning approaches, but their effectiveness depends heavily on how protein data are represented. Inspired by the success of language models in natural language processing, modern protein representation learning methods have emerged as powerful tools for capturing complex relationships between protein sequence, structure, and function. This study provides a comprehensive review of protein representation learning approaches, categorizing and evaluating them across multiple predictive tasks, including protein semantic similarity, functional annotation, drug target family classification, and mutation-induced changes in protein–protein binding affinity. The authors compare the strengths and limitations of these methods against traditional model-driven approaches, analyze the datasets and algorithms used, and discuss current challenges and future opportunities. Overall, the review highlights the transformative potential of representation learning in protein science and offers guidance for developing next-generation AI models for biological discovery.
1
1
11
741
An amazing review on ion channels—the building blocks of bioelectric signaling that power cellular communication, development, regeneration, and behavior in all living organisms, from bacteria to humans. #Bioelectricity #IonChannels #SystemsBiology Conduits of Life’s Spark: A Perspective on Ion Channel Research since the Birth of Neuron cell.com/neuron/fulltext/S08…
22
1,444
Overall structure of the Nav1.6 channel complex shown in side view (left) and cytoplasmic view (right). Glycan (sugar) moieties are depicted as stick representations, while the IFM motif within the intracellular III–IV linker is highlighted as spheres.
1
122
End the day by revisiting Cryo-EM structure of human voltage-gated sodium channel Nav1.6 Image source: PDB pnas.org/doi/10.1073/pnas.22…
1
2
12
844
On Connected Sublevel Sets in Deep Learning arxiv.org/abs/1901.07417
2
8
639
An artificial neural network integrated directly into computer memory enables real-time, high-accuracy reconstruction of human cortical function #Neuroscience #NeuralNetworks #NeuromorphicComputing #BrainInspiredAI #Neurotechnology #ComputationalNeuroscience #BioAI #BrainModeling #HumanCortex #CorticalNetworks #BrainSimulation #NeuroAI #DigitalBrains
2
5
23
1,173
This is highly concerning!!!! More than 18,000 questionable images found in antibody catalogues of 15 companies nature.com/articles/d41586-0…
1
1
6
591
Harmonic Torsional Diffusion for Protein-Ligand Flexible Docking arxiv.org/abs/2608.20366 1. Harmonic Torsion Modeling for Flexible Docking Harmony introduces a new diffusion docking framework that treats ligand and protein side-chain torsions as circular variables on a torus, rather than predicting angle updates in Euclidean space. Instead of directly regressing torsion angle changes, the model learns a Fourier (harmonic) potential for each rotatable bond and derives torsional score fields from its gradient. This explicitly captures periodicity, multimodal conformations, and rotameric side-chain behavior while remaining consistent with the physics of torsional motion. 2. Geometry-Aware Diffusion and Architecture The method incorporates an analytic frequency-dependent damping factor derived from variance-exploding (VE) diffusion on the torus. High-frequency Fourier modes naturally disappear at high noise levels and reappear during denoising, aligning the score model with the true forward diffusion process. Harmony jointly models translation, rotation, ligand torsions, and side-chain torsions on a product manifold using an SE(3)-equivariant e3nn graph neural network, making the harmonic torsion head a drop-in replacement for existing flexible docking architectures. 3. Improved Docking Accuracy and Physical Realism On PDBBind benchmarks with flexible pocket docking, Harmony substantially improves ligand pose prediction and pocket reconstruction accuracy compared with recent diffusion-based flexible docking methods. It also achieves better physical validity on PoseBusters tests, including gains in steric correctness, ligand strain, and protein–ligand interactions. Ablation studies show that harmonic torsion modeling, analytic frequency damping, and jointly learning ligand and side-chain torsions are all important contributors to performance. Take-Home Message Harmony's key innovation is replacing conventional torsion regression with a harmonic, geometry-aware representation of torsional score fields on the torus. By explicitly modeling periodicity and matching the mathematics of diffusion on circular variables, it achieves more accurate, physically realistic, and computationally efficient flexible protein–ligand docking. #AIDrugDiscovery #ProteinLigandDocking #DiffusionModels #GeometricDeepLearning #SE3Equivariance #GraphNeuralNetworks #ProteinEngineering #BioAI #FoundationModels #FourierAnalysis #HarmonicModeling #TorsionModeling #PDBBind #PoseBusters #Biophysics #NeuralNetworks
1
8
559
From Prompt to Drug: Toward Pharmaceutical Superintelligence pubs.acs.org/acscii/article/…
4
14
664
Reading World Models this weekend from @SchmidhuberAI lab arxiv.org/pdf/1803.10122
1
2
10
602
Simulating is not always understanding: When model complexity obscures biology arxiv.org/abs/2608.06998 1. Complexity Does Not Guarantee Understanding Computational models in cell biology range from simple systems with a few parameters to whole-cell simulations containing thousands of molecular components. However, a model advances understanding not through its size or complexity, but by generating novel predictions, uncovering unexpected biological relationships, or exposing missing mechanisms when it fails. 2. The Importance of Parameter Constraints The key determinant of a model's usefulness is the balance between the number of free parameters and the experimental data available to constrain them. Models built on a small set of well-grounded rules can reveal emergent behaviors and self-organization, whereas models with too many unconstrained parameters become difficult to interpret, validate, or falsify. 3. Toward More Interpretable Biological Models Rather than continuously increasing model complexity, the field should prioritize systematic comparisons between simple and complex models, dynamical analyses, and hierarchical modeling frameworks. Such approaches can better explain how cellular behaviors emerge from underlying biological processes while maintaining interpretability and predictive power. Take-home message: Understanding biological systems depends less on building the most detailed model and more on creating models that are interpretable, experimentally constrained, and capable of revealing underlying mechanisms.
2
12
57
2,813
A Generative Neuro-Symbolic AI for Protein Sequence Design advanced.onlinelibrary.wiley… EffieDes introduces a generative neuro-symbolic approach to protein inverse folding that moves beyond traditional auto-regressive sequence generation. Instead of predicting residues one at a time, the framework learns a backbone-conditioned probabilistic landscape and performs global optimization to identify protein sequences that best satisfy structural and functional objectives. This strategy enables more effective modeling of long-range residue interactions and avoids the limitations of greedy sequential sampling. 1. Interpretable Fitness Landscapes and Symbolic Optimization The framework consists of two components. First, EffieNN predicts pairwise residue interaction scores and constructs an explicit Potts-model representation of the protein fitness landscape. Second, an automated reasoning solver searches this landscape to identify optimal sequences while enforcing hard design constraints. By using a modified pseudo-likelihood objective (E-PLL), EffieNN better captures strongly unfavorable interactions that are critical for robust constrained design. The resulting fitness landscape is fully interpretable and can be optimized using exact or approximate combinatorial solvers, including scalability-enhancing methods such as LR-BCD. 2. Zero-Shot Constraint Handling and Exploration of Novel Sequence Space A major advantage of EffieDes is its ability to incorporate complex constraints without retraining the neural network. Constraints such as amino-acid composition limits, symmetry requirements, residue tying, and multi-state objectives are directly imposed during symbolic optimization. This enables zero-shot constrained protein design and facilitates exploration of highly unusual regions of sequence space. The framework successfully redesigned symmetric protein folds using extremely restricted amino-acid alphabets (as few as five residue types), generating sequences that retained strong folding potential despite being far outside the natural protein distribution. 3. Superior Performance in Protein Engineering Applications EffieDes outperformed Rosetta and several deep-learning inverse-folding methods on native sequence recovery benchmarks. In experimental studies, it demonstrated exceptional performance in multi-state interface design, converting bacterial microcompartment shell proteins into selective hetero-hexamers with an 86% success rate compared with 20% for ProteinMPNN. The framework also showed promise in de novo therapeutic protein engineering: an EffieDes-designed nanobody targeting the SARS-CoV-2 Omicron XBB.1.16 variant achieved nanomolar affinity, blocked ACE2 binding, and displayed variant selectivity, whereas none of the tested ProteinMPNN-designed candidates bound the target. Take-Home Message EffieDes bridges deep learning and symbolic reasoning to create an interpretable, globally optimizable protein-design framework. By enabling powerful zero-shot constraint handling and outperforming existing methods in challenging experimental design tasks, it represents a significant step toward programmable protein engineering in complex, low-data, and highly constrained biological settings. #ProteinDesign #GenerativeAI #NeuroSymbolicAI #ProteinEngineering #BioAI #DrugDiscovery
1
1
14
1,383
Temporal changes in metabolism guide oligodendrocyte precursor cell dynamics in aging and multiple sclerosis cell.com/neuron/fulltext/S08… Quick Summary: Aging disrupts the circadian clock regulator BMAL1 in oligodendrocyte precursor cells (OPCs), triggering metabolic dysfunction, cellular senescence, and impaired myelin repair. Remarkably, timed evening chronotherapy restores OPC function through SIRT2-dependent pathways. Similar BMAL1/SIRT2 disruptions are also observed in OPCs from multiple sclerosis patients, linking aging, circadian biology, and myelin regeneration. Takeaway: Circadian regulation of OPC metabolism may represent a promising therapeutic avenue for aging-related myelin decline and multiple sclerosis. #Aging #CircadianRhythm #BMAL1 #Myelin #OPCs #MultipleSclerosis #Neurodegeneration #Chronotherapy #Glia #CellBiology #NeuroAging
2
550
DeepGANnel: Synthesis of fully annotated single molecule patch-clamp data using generative adversarial networks pubmed.ncbi.nlm.nih.gov/3553… Automated analysis of single ion channel recordings using machine learning is limited by the lack of large, accurately labeled datasets. Because ion channel signals require sample-level annotations at very high temporal resolution, manually generating sufficient training data is impractical. This study addresses this challenge by developing a generative adversarial network (GAN)-based framework that can create unlimited amounts of labeled synthetic ion channel recordings from a small annotated dataset. The approach uses 2D convolutional neural networks (CNNs) to preserve the precise temporal relationship between raw electrophysiological recordings and their idealized channel-state labels. Unlike traditional simulation approaches, the method does not rely on predefined ion channel kinetics, noise assumptions, or hidden Markov models. Instead, the GAN learns the underlying structure directly from experimental recordings. The authors validated the generated synthetic data across five independent ion channel datasets, demonstrating that the artificial recordings closely resemble real data using dimensionality reduction techniques such as t-SNE and UMAP. This framework provides a scalable solution for creating training datasets for automated electrophysiology analysis and can potentially be extended to other biomedical time-series applications, including ECG signal analysis and nanopore sequencing. Take-home message: Generative AI can overcome the bottleneck of limited electrophysiology datasets by creating realistic, labeled synthetic recordings, enabling the development of robust machine learning tools for automated ion channel analysis and broader biological signal processing. #GenerativeAI #GANs #Electrophysiology #PatchClamp #IonChannels #BioAI #DigitalBiology
4
1
10
981
The Human Bindome: A Proteome-scale Atlas of Designed Binder Candidates biorxiv.org/content/10.64898… The Human Bindome: An Open Resource for Protein Binding The Human Bindome is a large-scale resource comprising 306,146 computationally designed protein binders targeting 8,296 human proteins. Each binder is accompanied by a fully specified amino acid sequence, a predicted binder–target complex structure, and confidence metrics. The project addresses long-standing limitations of antibody-based reagents, including high cost, inconsistent characterization, and lack of sequence transparency, by providing openly accessible and reproducible protein-binding reagents at proteome scale. Scalable Design Strategy and Proteome Coverage To generate the Bindome, the authors scaled the experimentally validated BindCraft pipeline using AlphaFold structural models and PAE-based domain segmentation. Proteins were divided into targetable domains based on structural confidence, compactness, and accessibility, while computational filters removed designs likely to clash with full-length protein structures. This strategy enabled efficient proteome-wide binder generation, resulting in coverage of 77.8% of the targetable human proteome and 40.9% of all human proteins, with an average of approximately 37 candidate binders per target protein. The resource also spans most major structural protein families, covering 78.2% of represented CATH superfamilies. Functional Targeting and Biological Relevance Analysis of binder epitopes revealed that designed binders preferentially target evolutionarily conserved and functionally important protein surfaces. Many binding sites overlap protein–protein interaction interfaces, ligand-binding regions, DNA-binding domains, post-translational modification sites, and known allosteric regulatory regions. The study demonstrates that binders can potentially perturb protein function by blocking critical interactions or regulatory mechanisms. Notably, 26% of annotated allosteric targets contain at least one binder overlapping a known allosteric site, highlighting the potential utility of these reagents for mechanistic biology, target validation, and therapeutic discovery. Structural Insights, Accessibility, and AI Applications The designed binder interfaces resemble native protein–protein interactions in size and composition while exhibiting largely novel interaction geometries. Although many binder scaffolds resemble known protein folds, their interfaces are distinct from naturally occurring complexes. To maximize usability, the entire dataset is distributed through a public web platform, a standardized 3D-Beacons API, and an MCP server that enables natural-language querying through AI agents. In addition, the authors provide leakage-controlled training, validation, and test splits comprising more than 248,000 examples, creating one of the largest resources available for developing and benchmarking next-generation machine learning models for protein binder design. Take-Home Message The Human Bindome represents a major step toward an openly accessible, proteome-scale catalog of programmable protein binders. Beyond serving as a powerful resource for biological research, it provides a foundation for AI-driven protein engineering, enabling systematic exploration of protein function, interaction networks, allosteric regulation, and therapeutic target discovery across the human proteome. #HumanBindome #ProteinDesign #ProteinEngineering #AIforScience #AlphaFold #BindCraft #DrugDiscovery #SyntheticBiology
1
4
21
1,208
The unintended consequences of large language models as a labor-augmenting technology in science arxiv.org/abs/2607.17397 Summary: 1. LLMs can accelerate scientific research by assisting researchers across multiple stages of the research pipeline, increasing overall productivity. 2. The impact of LLMs on publication quality depends on how they are used: when they help identify promising projects, researchers may become more selective; when they mainly simplify writing and publishing, researchers may become less selective. 3. Faster research increases the opportunity cost of scientists’ time, encouraging researchers to move more quickly between projects rather than spending additional effort refining existing work. 4. Greater efficiency does not necessarily lead to deeper thinking or better science. While LLMs save time on routine tasks, they may also create incentives that reduce thoroughness and alter how research effort is allocated. #LLMs #ScientificResearch #ResearchProductivity #ScienceInnovation #ResearchMethods #AIinScience #AIResearch #DigitalScience
2
4
26
4,121