arXiv · cs.CVBuildable★ flagship
Find and describe every cell in a microscope image without anyone labeling a single one.
Analyzing microscope images usually means someone hand-labels thousands of cells to train a detector, then runs separate steps to segment each cell and describe its type — tedious and annotation-hungry. This method skips the labels entirely: it learns to reconstruct each image by routing pixels to a small set of sparse 'source' points through a coarse-to-fine pyramid, and in doing so it naturally discovers which pixels belong to which cell. Those pixel-to-source groupings become the instance masks, while each source's compressed code captures the cell's shape and morphology, giving you both segmentation and a phenotype fingerprint at once. Because it's unsupervised, it works across different cell shapes and imaging setups without retraining on new annotations. It matters because manual labeling is the main bottleneck in scaling up biological image analysis.
Technical view
The method performs unsupervised cell instance segmentation and phenotypic representation by reconstructing each image with a coarse-to-fine 'generative routing pyramid' that associates pixels with spatially sparse latent sources. The pixel-to-latent assignments directly yield instance masks, while the per-source latents encode morphology usable for phenotypic classification — unifying segmentation and representation in one generative pass rather than the usual detect-then-featurize pipeline. Reported results show competitive instance-segmentation performance across diverse cell morphologies and imaging modalities plus phenotype/gene-related analysis, all without manual annotations. Practitioners could apply it to unlabeled microscopy corpora to bootstrap masks and morphology embeddings, or adapt the sparse-routing reconstruction objective as an annotation-free pretraining stage.
arXiv · q-bio.BMConceptual★ flagship
Read a protein's shape straight from blurry microscope snapshots—skipping the usual 3D map.
Cryo-electron microscopy freezes millions of copies of a protein and photographs them from random angles, but each photo is a noisy, flattened shadow rather than a clear 3D picture. Normally scientists first stitch these shadows into a fuzzy 3D density map and then guess where the atoms sit; this work skips that middle step and fits the protein's atomic skeleton to the raw photos directly. They start with a rough template of the backbone and gently bend and twist it—like reshaping a wire model—until the shadows it would cast match the real photos. Because proteins wiggle between different shapes to do their jobs, being able to catch those shape changes matters for understanding how they work and how drugs might target them. Doing it in one direct step could make the reconstruction cleaner and better at capturing motion.
Technical view
The method poses backbone recovery as indirect shape matching: an atomic point-cloud template is deformed so its simulated tomographic projections (through the cryo-EM imaging operator) match observed single-particle images, bypassing intermediate 3D potential-map reconstruction. The deformation is driven by a gradient flow on a Lie group, derived first in a general geometric setting then specialized to the SPA forward model. On synthetic data it recovers single- and multichain proteins and captures conformational transitions. A practitioner could extend it toward heterogeneous/continuous conformational analysis and, eventually, experimental data by plugging in realistic CTF and noise models into the projection operator.
arXiv · q-bio.QMBuildable
An AI that predicted premature birth almost too well — because it was secretly cheating.
Doctors can pick up tiny electrical signals from a pregnant woman's abdomen (called electrohysterography, or EHG) that reflect the uterus's muscle activity, and researchers have tried using these signals to predict preterm birth. The catch: many past studies let recordings from the same patient show up in both the 'training' and 'testing' data, which is like letting a student see the exam answers beforehand — it makes the AI look smarter than it is. This paper builds a fairer test where each patient's data stays entirely on one side, and adds a system that flags cases the model is unsure about instead of forcing a guess. The goal is a more honest, trustworthy tool for identifying real preterm-birth risk.
Technical view
The authors formalize segment-level vs. patient-independent (record-grouped) validation on the Term-Preterm EHG Database (300 records, 38 preterm) and benchmark a 92-feature elastic-net logistic model under record-grouped nested cross-validation, with preprocessing, Platt calibration, and conformal estimation strictly confined to training folds. They further implement class-conditional conformal selective prediction to abstain on low-confidence cases rather than force uncertain classifications. This establishes a leakage-proof baseline other EHG classifiers can be benchmarked against using the same evaluation protocol.
arXiv · eess.SPBuildable
Splitting a pregnant belly's electrical hum into layers to hunt for early-birth warning signs.
EHG signals from the abdomen carry a mix of overlapping electrical rhythms from the uterus, much like a song with multiple instruments playing at once. This study uses a technique called empirical mode decomposition to separate that mixed signal into simpler layers, then tests which layer (and which features extracted from it) best distinguishes women who deliver preterm from those who deliver at term. They also compare using expert-picked recording snippets versus plain fixed time windows, to see which gives more reliable results. The idea is to find a signal-processing recipe that's both accurate and doesn't accidentally cheat by mixing the same patient's data across training and testing.
Technical view
Using 26 recordings (13 preterm, 13 term) from the public TPEHGT dataset, the authors decompose EHG signals via empirical mode decomposition into intrinsic mode functions (IMFs) and extract 14 features per channel across 3 channels, comparing annotated intervals against non-overlapping 3-minute fixed windows. Nine classifiers are evaluated with repeated 5-fold recording-grouped cross-validation and recording-level aggregation to avoid the leakage problem noted in [1]. IMF1 (the highest-frequency component) gave the strongest mean classification performance, suggesting fast oscillatory content carries the most discriminative signal for term/preterm classification.
arXiv · q-bio.QMBuildable
A cell-map tool that doesn't squash crowded neighborhoods into looking empty.
When scientists visualize thousands of individual cells based on their gene activity, they use 2D maps (like UMAP) to see clusters of similar cell types. But these maps often distort how 'crowded' or 'sparse' different regions really are, which matters when you're trying to spot rare or in-between cell states. DMT-Dens is a new mapping method, built on a transformer-based neural network, that specifically preserves this density information alongside the usual neighborhood structure. It works by making sure that how tightly packed points are in the original high-dimensional data matches how tightly packed they appear in the final 2D picture. This gives biologists a more trustworthy visual to spot rare or transitional cell populations.
Technical view
DMT-Dens is a parametric manifold-visualization method using a latent-token Transformer encoder that combines rank-based manifold alignment with hard-pair aggregation for neighborhood preservation. Its key addition is a density-preservation loss based on the Pearson correlation between k-nearest-neighbor log-radius estimates computed in the original high-dimensional space versus the 2D embedding space. Benchmarks show strong density fidelity on biological single-cell datasets, making it a drop-in alternative to UMAP/t-SNE when density-aware interpretation (e.g., detecting rare or transitional populations) matters, and being parametric, it can embed new/unseen data without retraining.
arXiv · q-bio.QMBuildable
An AI that designs new molecules by exploring branching 'what-if' possibilities like a chess engine.
Designing new biomolecules — proteins, but also trickier targets like DNA and RNA — that bind to a specific partner is central to drugs and biotech, but there's far less training data for DNA/RNA than proteins. MCTH tackles this by using existing AI models that predict 3D shapes from sequences (and vice versa) as building blocks, then uses a search strategy borrowed from game-playing AI (Monte Carlo Tree Search, the technique behind AlphaGo) to explore many candidate designs and spend its computing budget on the most promising ones. It factors in how confident the underlying models are and whether multiple predictors agree, and can optionally steer designs toward specific physical properties. This offers a way to design new molecule pairs without needing to retrain the underlying AI models.
Technical view
MCTH (Monte Carlo Tree Hallucination) is an inference-only framework that frames all-atom sequence-structure co-design as uncertainty-aware planning: it treats pretrained folding and inverse-folding models as frozen black-box operators, generating 'hallucinated' candidate states, and uses Monte Carlo Tree Search to allocate a fixed inference budget across competing design trajectories. Node selection incorporates model confidence/uncertainty and cross-expert consensus/disagreement when multiple predictors are available, with optional biophysical constraints folded into the same decision loop. Because it requires no retraining, practitioners can plug in any folding/inverse-folding model pair and apply it to non-protein modalities like DNA/RNA where labeled complex data is scarce.
arXiv · q-bio.QMBuildable
A cell-clustering AI you can actually peek inside to see why it made each call.
When AI groups similar cells together from gene-expression data (single-cell RNA sequencing), it usually works like a black box — you get clusters but can't easily see the reasoning. scDNM-VAE is a new model inspired by how brain neurons process signals through branching dendrites, where each 'gate' has a clear direction, strength, and threshold for how it responds to the data. Because these gates are simple and explicit rather than buried in an opaque network, researchers can directly read off why a cell was assigned to a given cluster, without needing a separate explanation tool bolted on afterward. Tested on immune, brain, heart, and blood-stem-cell data, it holds its own against standard methods while being more transparent.
Technical view
scDNM-VAE pairs a variational autoencoder with a dendritic-neuron-inspired clustering head where cluster assignments are governed by learnable signed synaptic weights and thresholds: weight sign sets gate response direction, magnitude sets steepness, and the weight-threshold pair sets the transition location in latent space. This makes the trained clustering function directly inspectable without post-hoc explainability methods (e.g., SHAP/LIME analogs). It's benchmarked against scVI+KMeans and an MLP-DEC ablation across four datasets spanning immune, cortical, cardiac, and hematopoietic cells, offering a template for building interpretable-by-construction deep clustering models in other domains.
arXiv · nlin.AOConceptual
Reading the 'wave shapes' in brain or network rhythms like a fingerprint of their spatial pattern.
Many systems — from groups of neurons firing together to engineered oscillator networks — form visible spatial patterns as they pulse in sync or drift out of sync, and these patterns can shift suddenly and briefly (transient dynamics), which is hard to catch. This paper introduces a way to describe those patterns by looking at the timing (phase) of oscillations at nearby points and ranking their relative order, rather than looking at how strong the signal is (amplitude). From this ranking, they compute a single number — a kind of 'diversity score' — that rises when many different spatial patterns are present and can flag the moment a system briefly switches behavior. It's a general lens for studying rhythmic, spatially-spread-out systems, from brains to power grids.
Technical view
The method extends ordinal-pattern symbolic analysis to the spatial domain, operating directly on instantaneous phase rather than amplitude, with extra symbols added to handle near-equal phase values across neighboring points. This yields a symbolic representation encoding local spatial ordering that simultaneously captures phase gradients and synchronized clusters, from which a spatial permutation entropy is defined to quantify pattern diversity at each timepoint. The entropy time series enables detection of transient dynamics and regime shifts in oscillatory systems, giving practitioners a computationally light, model-agnostic diagnostic applicable to neural recordings or engineered oscillator networks.
arXiv · q-bio.PEConceptual
Hitting 'pause' lets rock-paper-scissors species dodge the random extinctions that would otherwise wipe them out.
In ecosystems where species compete in a rock-paper-scissors style loop (each type beats one and loses to another), random population swings can accidentally wipe out a type entirely, collapsing the diversity even though no species is actually superior. Scientists knew that physical space — separate patches acting as refuges — can protect against this. This paper asks whether something similar can happen in time instead of space: specifically, whether organisms going dormant (like seeds lying inactive in soil, or bacteria entering a resting state) can act as a 'time refuge' that keeps lineages alive through unlucky stretches. Using a mathematical population model, they show dormancy indeed prevents this random collapse, offering a new explanation for how competitive diversity persists even in well-mixed, unstructured populations.
Technical view
The authors build a discrete-time Wright-Fisher population-genetic model that combines generalized seed banks with frequency-dependent (non-transitive, rock-paper-scissors-like) interactions, where an individual's type can be inherited from potential parents sampled across multiple past generations rather than only the immediately preceding one. This dormancy mechanism acts analogously to spatial structure, buffering lineages against interaction-driven stochastic fluctuations that would otherwise drive the system to fixation/extinction in a standard well-mixed model. The framework provides a tractable population-genetics tool for studying how temporal refuges (dormancy, seed banks) stabilize coexistence, applicable to microbial, plant seed-bank, or other systems with dormant life stages.
arXiv · math.APConceptual
A math proof settling which of two competing species wins the turf war along their border.
Imagine two species competing fiercely for the same space, spreading out and bumping into each other along a moving boundary — like two colors of mold racing across a petri dish. Mathematically, this is modeled with equations (Lotka-Volterra competition-diffusion) that predict a wave-like front between the two territories, and the key question is which species pushes the front forward and claims more ground. This paper proves, in full generality for the case where both species have identical competitive strength, that the species which spreads out (diffuses) faster always wins and expands its territory — except in the special case where both spread at exactly the same rate, where neither wins. It's a clean, complete answer to a question ecologists and mathematicians have long puzzled over.
Technical view
For the symmetric two-species Lotka-Volterra competition-diffusion system under strong competition (competition intensity >1), the authors fully characterize the sign of the unique bistable traveling front's speed: for every diffusion ratio d≠1, the front always expands the territory of the faster-diffusing species, with zero speed exactly at d=1. The key technical lemma is that no monotone standing front can exist when the two diffusion rates differ, which combined with continuity of wave speed in the model parameters yields the sign result; they also establish smooth dependence of the wave speed and front profile on parameters. This closes an open question in reaction-diffusion theory and gives a rigorous basis for predicting invasion outcomes in reaction-diffusion competition models from diffusion rates alone.
arXiv · math.PRConceptual
In simulated brain circuits, the order tiny signals arrive—not just their size—can flip whether a neuron fires.
This is a math study of simplified 'spiking' neuron networks, where neurons fire once their input crosses a threshold and then reset. Neuroscientists often simplify fast synaptic signals by shrinking their timing down to an instant, assuming only the total amount of excitation and inhibition matters, not the precise sequence they arrive in. The authors build two toy networks where the excitatory and inhibitory nudges shrink to the same instantaneous size in the limit, but arrive in opposite order — excitation-then-inhibition versus inhibition-then-excitation — and show a target neuron fires in one case but not the other. It matters because it exposes a hidden flaw in standard mathematical shortcuts used to model fast brain circuits: they can quietly give the wrong answer about whether a neuron actually fires.
Technical view
Uses a causal event-driven protocol with clamped refractoriness and smooth positive-delay kernels to construct two families of signed synaptic measures that converge weakly to δ₀ while their microscopic arrival order is reversed. The target neuron's firing condition, x+a−b<θ≤x+a, depends on this order rather than on the limiting measure alone, and the effect is shown to be robust to perturbations of state, pulse mass, and drift, persisting on sparse Dale-compatible random block graphs with q_N→∞, q_N/N→0. This implies that naive fast-synapse (instantaneous) limits of E/I spiking network models can be discontinuous or ill-posed, which matters for anyone deriving mean-field or diffusion approximations of spiking networks.
arXiv · cs.LGBuildable
Training an AI on real cell-biology experiments—not textbooks—makes it reason like a biologist.
PertMind trains a large language model using actual lab measurements of how genes respond when a cell is perturbed, say by knocking out a gene or applying a drug. Instead of relying on expensive human-written explanations, it treats the measured outcomes as a reward signal in reinforcement learning, like a game score telling the model whether its prediction was right. The model starts with a supervised warm-up on trusted example reasoning, then improves through feedback scored at the level of individual genes, whole pathways, and answer formatting. Because it learned by predicting how perturbations play out, it also got better — without any extra training — at related tasks like figuring out what perturbation caused an outcome or picking the most promising experiments, suggesting it absorbed real biological intuition rather than memorized facts.
Technical view
PertMind combines supervised initialization on trusted reasoning trajectories with RL using multi-level rewards (gene-, pathway-, and format-level) computed directly from cellular perturbation atlas measurements, training only on forward perturbation-response prediction. It reports improved generalization to unseen cellular contexts plus zero-shot transfer to reverse perturbation identification, double-perturbation reasoning, phenotypic-screen prioritization, and biological-process interpretation without task-specific fine-tuning. This is a reusable template for converting large biological measurement atlases into RL environments with computable, non-human-curated rewards for post-training domain LLMs.
arXiv · q-bio.QMConceptual
A math toolkit tells therapists exactly which symptom to target, when, and how hard to push.
Psychologists increasingly model mental-health symptoms — sleep trouble, sadness, anxiety — as a network where each symptom can influence others over time. Knowing that symptom A tends to predict symptom B later doesn't tell a clinician what to actually do: how much to shift A, or whether nudging it will meaningfully change B down the line. tSymPerturb turns these descriptive prediction networks into an actual playbook, defining ways to simulate 'turning down' a symptom, blocking the pathway between two specific symptoms, and testing strategies like dosage, combining interventions, and choosing the best order to apply them. A core equation spells out exactly how much a later symptom will shift based on how much an earlier one was changed, turning a correlational map into something closer to a testable treatment-design tool.
Technical view
tSymPerturb extends the SymPerturb formalism to cross-lagged panel networks (CLPNs), separating source-state operators (virtual knockout/knockdown), transition operators (directed-edge or source-node communication blocking), and strategy procedures (dosage-response, combination analysis, sequence optimization). For a two-wave linear CLPN, the propagation identity Δμ₂ = B(μ₁ − μ₁*) makes explicit how a perturbation at wave 1 propagates through transition matrix B to wave-2 outcomes, and the paper derives falsifiable predictions — e.g., exact linear dose-response under a fixed intervention — that researchers can test against real panel data. This gives anyone with longitudinal CLPN data a route from correlational symptom networks to simulate-able intervention design.
arXiv · q-bio.NCBuildable
How an AI avoids forgetting old skills quietly determines which of its internal memories drift over time.
Brains and AI systems both face the same puzzle: how do you learn new things without erasing what you already know, and does the trick you use to protect old memories change how those memories subtly shift over time? Researchers trained artificial networks — both image-classifiers and networks handling sequences of cognitive tasks — using different anti-forgetting tricks, most notably 'replay,' where the system periodically re-practices old material while learning new material. They tracked how the network's internal representation of the same fixed test inputs changed session after session, like repeated brain scans of the same thought. Replay kept performance intact, but the internal representations still drifted, and not randomly: deeper, more detailed processing stages wandered the most while coarse category information stayed stable — mirroring drift patterns seen in real animal brains, which suggests it's a natural byproduct of balancing stability and new learning.
Technical view
CNNs were trained on sequential image-classification tasks and RNNs on sequences of cognitive tasks under various continual-learning regularizers, with experience replay tracked in detail, while monitoring drift in fixed-probe representations across intervening tasks. Replay prevented catastrophic forgetting in both architectures, but representational drift still accumulated monotonically with the number of intervening tasks and was structured: later visual-processing layers and RNN temporal tuning drifted more than coarse class organization or task-relevant readout directions. This gives a testable, mechanistic link between specific continual-learning algorithms and drift signatures, letting neuroscientists compare recorded drift statistics in animals against model predictions to infer which mechanism better explains cortical drift.
arXiv · cs.LGBuildable
A brain-computer interface keeps working day after day by learning geometry, not just raw brain data.
Motor imagery brain-computer interfaces let a device guess what movement someone is imagining just from their brainwaves, potentially letting paralyzed patients control a wheelchair or robotic arm by thought. Two big problems plague real clinical use: brain signals look different day to day, and the system must adapt on the fly without stopping to retrain. MRieHy tackles this by first mathematically aligning each day's brain-signal patterns onto a shared reference using Riemannian geometry, a technique that treats the signal's covariance structure like points on a curved surface, then builds two web-like 'hypergraphs' connecting similar patterns — one from this geometric similarity, one from learned features — so the system can borrow strength from many related samples at once. Combining both lets the interface keep recognizing imagined movements accurately even as raw signals drift across days, tackling a major obstacle to everyday BCI use.
Technical view
MRieHy computes Riemannian means of EEG covariance matrices across cross-day training sessions to align multi-day distributions, then builds two complementary hypergraphs — one over covariance matrices using Riemannian distance, another over deep feature embeddings — to capture higher-order sample relationships during online test-time adaptation. This combines Riemannian-geometry domain alignment, standard in EEG transfer learning, with hypergraph message passing for higher-order (beyond pairwise) relations, applied specifically to the streaming/online setting rather than offline recalibration. Practitioners building MI-BCI pipelines could adopt the dual-hypergraph fusion as a drop-in module for test-time adaptation atop existing Riemannian-alignment baselines.
arXiv · q-bio.NCConceptual
A control-theory equation tries to pin down exactly which brain circuit makes things conscious.
Global Workspace Theory is a popular idea in consciousness science: the brain becomes aware of something when that information gets broadcast widely from a central hub to the rest of the brain. The problem is nobody has a precise, checkable definition of what counts as that hub. This paper borrows tools from control theory — the math engineers use to analyze how systems like thermostats or autopilots respond to and influence their surroundings — to define the hub as a subnetwork that can be driven by the rest of the brain, can in turn drive the rest of the brain, and has internal 'modes' linking the two directions in a distinctive way. By turning 'receives input,' 'sends output,' and 'transforms information' into precise mathematical properties, the authors create a testable signature that could, in principle, be checked against real or simulated brain circuits to see whether a candidate region really behaves like a global workspace.
Technical view
The Global Mediation Workspace (GMW) formalizes a candidate global-workspace subnetwork as an open dynamical system embedded in a larger network, using reachability (a Gramian capturing how external inputs drive the subnetwork), observability (how subnetwork states affect the rest of the network), and a boundary Hankel operator that identifies the internal modes coupling input-driven and output-driving dynamics. This gives a quantitative, control-theoretic alternative to informal 'broadcasting' language in Global Workspace Theory. It could be applied to whole-brain or large-scale RNN models by computing Gramians/Hankel singular values from simulated or empirical connectivity to test which subnetworks satisfy the GMW signature.
arXiv · q-bio.QMBuildable
AI reads scattered heart sensors and pinpoints the exact scarred tissue causing dangerous heartbeats.
When doctors treat irregular heartbeats with a procedure called ablation, they need to find the exact patch of heart tissue causing the problem, but they can only take readings from a limited, sparse set of points inside the heart. This project trains a graph neural network — an AI good at reasoning over networks of connected points — on realistic simulated heart-signal data to spot suspicious regions, like scarred tissue, unusually fast-firing areas, or overly excitable tissue, all of which can trigger a dangerous rhythm called premature ventricular complexes. The model detected these regions with very high accuracy on simulated flat hearts, and with just a little extra fine-tuning it generalized to curved, more realistic heart shapes — a promising step toward guiding cardiologists to the right ablation target using fewer invasive measurements.
Technical view
The authors trained a GNN on synthetic electrogram data over 2D flat surfaces to classify localized regions of interest for PVC ablation, achieving average precision of 0.96 (fibrosis), 0.97 (rapid depolarization), and 0.95 (high excitability) from sparse intracardiac sampling. The model transfers to curved 2D surfaces via few-shot fine-tuning, indicating the learned graph representation generalizes beyond flat training geometry — a step toward geometry-agnostic clinical deployment. Real clinical use would still require validation on real patient electrograms and full 3D cardiac anatomy rather than synthetic flat-surface data.
arXiv · q-bio.QMBuildable
An AI agent invents, tests, and revises its own scientific theories to design better algorithms.
This project asks: what if you gave an AI agent the actual scientific method — form a hypothesis, build it, test it, learn from results, repeat — and set it loose on inventing better algorithms? 'The Little Scientist' has a 'Scientist' AI agent work inside a testing environment that runs its code and reports detailed feedback on each case, much like a lab assistant handing back experiment results. When the Scientist gets stuck improving its own ideas, a second AI called the 'Kuhn agent,' named after the philosopher who coined 'paradigm shift,' steps in and throws it a wildly different idea borrowed from an unrelated field, forcing it to explore a totally different approach instead of endlessly tweaking the same one. The idea is that automated discovery needs deliberate disruption, not just iteration, to escape dead ends — mirroring how real scientific breakthroughs often come from outside conventional thinking.
Technical view
The Little Scientist implements an iterative loop where a Scientist LLM agent proposes and implements algorithm designs, evaluated by a benchmark environment returning structured per-instance diagnostics; on detecting a performance plateau, a Kuhn agent injects a cross-disciplinary 'paradigm-shifting' conjecture to redirect search away from the local optimum in the LLM's latent solution space. This is effectively an LLM-agent-driven automated algorithm design system with a built-in exploration/exploitation controller triggered by stagnation detection, demonstrated on two problems requiring different reasoning modes. Practitioners building LLM-agent AutoML or algorithm-search pipelines could adopt the plateau-detection plus cross-domain-analogy-injection pattern as a general escape-local-optima mechanism.
arXiv · cs.LGBuildable
AI reads the 3D shape of your eye's optic nerve to trace your ancestry.
Researchers studied a small, tight-knit community on Norfolk Island in the Pacific, many of whom descend from the Bounty mutineers, by looking at photos of the back of their eyes instead of their DNA. They took two-angle (stereo) photos of the optic nerve head — the spot where the eye connects to the brain — and computationally rebuilt its 3D shape, since that shape is partly inherited. A deep neural network then broke this shape down into layers of detailed features, and the most genetically telling features were used to sort people into ancestry-related clusters. The point is to show that a cheap eye photo can reveal population ancestry patterns that normally require expensive genetic testing.
Technical view
The pipeline reconstructs 3D optic nerve head (ONH) morphology from stereo fundus photographs via multi-scale stereo matching, then applies a self-taught deep learning model to extract hierarchical shape features at multiple scales. Features are ranked by discriminant power for distinguishing Bounty-descendant vs. non-descendant subpopulations within 781 Norfolk Island individuals, with performance validated via stratified cross-validation. Selected features feed hierarchical k-means-style clustering (k=2–7) to estimate admixture-like membership fractions, offering a low-cost phenotypic proxy for genetic population structure in imaging-based epidemiology.
arXiv · q-bio.NCConceptual
When brain regions talk back, a strange trick where the follower predicts the leader can break down.
In the brain, two connected regions can sync their rhythms, and oddly, sometimes the 'receiving' region seems to anticipate the 'sending' region rather than lag behind it — a bit like a dance partner predicting your next move. Scientists usually study this using a one-way connection, but real brain areas send signals back and forth. This paper adds that return signal (excitatory feedback) to computer models of two connected neuron populations and watches what happens to the anticipation effect and to a related phenomenon where the system can flip between two synchronization states. Understanding this helps explain confusing timing patterns seen in real brain recordings, where it's not obvious which region is 'leading.'
Technical view
The study extends unidirectional cortical-population models exhibiting anticipated synchronization (AS, negative phase lag from receiver having faster intrinsic dynamics) by adding excitatory feedback from receiver to sender, forming a bidirectional motif. Using coupled neural-mass-type oscillator models, the authors characterize how feedback strength modulates the existence and stability of AS versus delayed synchronization (DS), as well as the bistable regime between them. Results clarify how bidirectional cortical coupling — closer to physiological reality than the previously studied unidirectional case — shapes phase-lag statistics observed in electrophysiology, informing interpretation of lead-lag relationships in real inter-areal recordings.
arXiv · q-bio.TOBuildable
Scientists dropped a crash-test dummy with live brain cells inside its head to watch neurons react to impact.
To understand what actually happens to brain cells during a head injury, researchers built a crash-test-dummy-like full-body model with real, living neurons embedded inside its head in small dishes. They dropped the dummy from a seated position at different angles to simulate falls, then measured the forces on the head with accelerometers while simultaneously tracking how the living cells responded to the jolt. A separate computer model of human muscles and bones was used to double-check the fall's physics. The goal is to directly connect the mechanical violence of an impact to the biological damage it causes at the cellular level, which could improve helmet design and injury thresholds.
Technical view
The framework integrates a commercial anthropomorphic surrogate with three stacked Petri dishes of live SH-SY5Y neuroblastoma cells inside the head, instrumented with six accelerometers (three head-surface, three in-series with the cell stacks) to capture impact kinematics during controlled seated falls at 30°, 60°, and 90° release angles. An OpenSim musculoskeletal model runs in parallel to reproduce fall biomechanics, enabling correlation of measured head acceleration/deformation with observed cellular response. This links macro-scale biomechanical impact metrics directly to micro-scale cellular outcomes, offering a validation platform for injury-threshold and protective-equipment research.
arXiv · q-bio.PEConceptual
Math shows why chemo that looks like it 'cures' a tumor on paper can still let it come back.
Standard cancer-treatment math assumes that if you give a high enough dose of chemotherapy, the tumor's cell count smoothly goes to exactly zero — a clean cure. This paper argues that's an illusion caused by ignoring randomness: real cell populations are small, discrete, and noisy, especially in tiny hidden pockets of tumor cells shielded from the immune system. Using advanced physics techniques for modeling random fluctuations (originally built for other 'noisy population' problems), the authors show these tiny surviving clusters have a genuinely nonzero chance of regrowing even after treatment that looks perfect on average. This matters because it offers a mathematical reason why cancers relapse even after seemingly successful chemo, and could inform how doses are scheduled.
Technical view
The authors build a nonequilibrium stochastic PK-PD field theory coupling a two-compartment pharmacokinetic model to a stochastic tumor-immune sector, formalized in the Doi-Peliti operator formalism and mapped to multiplicative Langevin equations via the Martin-Siggia-Rose/Janssen-De Dominicis path-integral approach. In immune-depleted sanctuary sites the dynamics reduce to a time-dependent Feller diffusion process, and the corresponding Fokker-Planck equation yields a closed-form survival functional showing that demographic noise gives micro-clusters a strictly positive relapse probability even under deterministic-cure-predicting high-dose bolus chemotherapy. This provides an analytical, testable framework for reassessing dosing strategies (e.g., metronomic vs. bolus) against fluctuation-driven relapse risk.
arXiv · cond-mat.softConceptual
Growing tissue isn't just solid or liquid — push it fast enough and it does something entirely new.
Living tissues like tumors or biofilms grow, and as they grow they also respond to stress like a mix of a solid (springy, elastic) and a liquid (flowing, viscous) — think of Silly Putty. This paper asks what happens when the speed of growth starts to compete with how quickly the material relaxes stress internally. Using a simple test case — a growing elastic beam — the authors find that when growth and relaxation happen at similar speeds, the material doesn't just gradually shift between solid-like and liquid-like behavior; it undergoes a genuinely new kind of transition with its own distinct dynamics. This reshapes how we should think about the mechanics of anything that grows while also being squishy, from tumors to bacterial colonies.
Technical view
The paper analyzes proliferating viscoelastic matter where growth rate g competes with the material's viscoelastic relaxation time τ, showing the combined dimensionless parameter gτ governs a qualitative transition rather than a smooth interpolation between the purely viscous (gτ→0) and purely elastic (gτ→∞) limits. Using a growing elastic beam as the canonical test system, they identify new dynamical regimes emerging at intermediate gτ that are absent from either limiting theory. This establishes a general theoretical lens — applicable to biofilms, tumors, and other proliferating soft matter — for predicting when growth-induced mechanical instabilities (e.g., buckling, morphogenesis) will deviate from standard elastic or viscous growth models.
arXiv · q-bio.PEConceptual
A review of the math tools conservationists use to save the whole 'tree of life,' not just species counts.
When deciding which species to prioritize for conservation, just counting species can miss the bigger picture — losing one weird, evolutionarily unique species (like a platypus) is a bigger loss to life's diversity than losing one of many similar frog species. Since the early 1990s, scientists have built mathematical tools that measure diversity using the 'tree of life,' the branching family tree connecting all species, so that older, more distinct branches count for more. This chapter walks through how those tools evolved, especially a widely used one called phylogenetic diversity and its descendants (like EDGE and EDGE2), which rank individual species by how much unique evolutionary history they'd take with them if they went extinct. It matters because it gives conservationists a more principled way to decide where limited funding and effort should go.
Technical view
The chapter reviews the methodological lineage of phylogenetic diversity (PD) metrics, starting from Faith's PD (sum of edge lengths in a rooted phylogenetic subtree) through species-specific prioritization indices like EDGE (Evolutionarily Distinct and Globally Endangered) and its refinement EDGE2, which quantifies expected marginal PD loss per species under extinction risk. It surveys the mathematical formalization and extensions of this framework for use in conservation prioritization, providing a technical primer for practitioners implementing PD-based ranking in biodiversity assessment or reserve-design software.
arXiv · q-bio.NCConceptual
A new filing system lets AI research assistants share and organize scientific knowledge across an entire team.
AI assistants that help with science increasingly rely on memory systems — knowledge graphs — to keep track of files, ideas, and results over long projects. The problem is that most of these systems are built around one person's personal organizing habits, making it hard for a whole team to share and reorganize that knowledge together. Valhalla is a proposed framework that replaces the usual flat, tangled web of notes with organized layers: separate levels for raw files, extracted resources, defined concepts (entities), the relationships between them, and the overall graph. This layered structure is meant to make long-term scientific knowledge easier to trust, share, and rebuild as a research team's understanding evolves.
Technical view
Valhalla introduces a five-layer File-Resource-Entity-Relationship-Graph (FREG) architecture as an alternative to conventional flat, node-centric knowledge graphs used for LLM-agent long-term memory in scientific workflows. Files and Resources preserve source identity and provenance, while Entities and Relationships are built as stable semantic abstractions on top, layered under a governing Graph service — aiming to decouple knowledge structure from any single user's ad hoc organizational scheme and enable cross-user sharing, integration, and reorganization. This targets a concrete pain point in multi-agent/multi-user RAG and knowledge-management systems: provenance-preserving, reorganizable shared memory, relevant to anyone building persistent knowledge infrastructure for LLM research agents.
arXiv · q-bio.NCBuildable
A single, well-timed electrical pulse can push a network of neurons into or out of sync — if you hit it right.
Groups of neurons often fire in rhythmic waves, and scientists want to know how a brief external nudge — like a pulse of current, similar to what a stimulation device might deliver — changes that rhythm. The usual tool for this, the phase response curve, only tracks how the pulse shifts the timing of the rhythm, but it misses how the pulse also changes the strength or intensity of the synchronized activity. This paper studies a simulated network of excitatory and inhibitory neurons and tracks both the timing shift and the intensity shift caused by pulses delivered at different points in the rhythm. They find that the exact same pulse can make the network more synchronized or less synchronized purely depending on when in the cycle it arrives, which matters a lot for designing brain-stimulation therapies that aim to calm or boost rhythmic brain activity.
Technical view
The authors simulate a balanced excitatory-inhibitory network of exponential integrate-and-fire (EIF) neurons receiving phase-targeted transient current pulses, and jointly compute the network phase response curve (nPRC), a novel network amplitude response curve (nARC), and resulting changes in population synchrony as functions of stimulus timing. They demonstrate that identical pulses can enhance or suppress synchronization depending solely on pulse phase, showing phase-resetting theory alone (standard PRC analysis) is insufficient to predict collective network response — amplitude effects are equally causal. This nPRC/nARC framework gives a quantitative, testable basis for designing closed-loop or phase-locked neurostimulation protocols aimed at modulating pathological or therapeutic network synchrony (e.g., in epilepsy or Parkinsonian oscillations).
arXiv · q-bio.NCConceptual
How long it takes one neuron to nudge another secretly tunes whether brain rhythms wobble or snap back.
Brain cells talk to each other through synapses, and there's always a tiny delay before a signal from one neuron actually affects the next. This study looks at brainwave-like rhythms produced when excitable 'go' neurons and calming 'stop' neurons volley signals back and forth in a fast, well-known rhythm pattern seen in the cortex. The researchers poked these simulated networks with brief jolts and measured two things: whether the rhythm's timing got shifted (like a beat dropping early or late) and whether its strength changed. They found that the delay length itself — not just the network's wiring — determines how the whole population recovers from a disturbance, which matters for understanding how the brain keeps rhythms stable or lets them shift, something linked to attention, memory, and disorders like epilepsy.
Technical view
Using a conductance-based spiking network in the PING (pyramidal-interneuron gamma) regime, the authors systematically vary synaptic delay and apply brief perturbations to excitatory (E), inhibitory (I), or combined populations, then compute network phase response curves (nPRCs) and amplitude response curves (nARCs). This extends single-neuron PRC theory to a network observable, showing delay reshapes both timing-reset and amplitude-recovery dynamics of the collective oscillation, not just the E/I balance typically emphasized. Practitioners modeling gamma oscillations or building delay-coupled E-I mean-field models can use nPRC/nARC as a diagnostic to predict entrainment and desynchronization thresholds under stimulation (e.g., for closed-loop neurostimulation design).
arXiv · eess.SPBuildable
One math trick reads your heartbeat and breathing from almost any wearable sensor, cutting through noise automatically.
Doctors and wellness apps need to track heart rate and breathing rate from all sorts of sensors — chest straps, fingertip clips, hospital monitors — but each sensor type usually needs its own custom software to filter noise and find the rhythm. CORAL is a new general-purpose method that turns any repeating, wave-like signal into a special 2D picture (they call it a 'correloform') that highlights how periodic the signal is over time, using a classic statistical tool called autocorrelation. Because this picture-making process is built from math rather than tricks tailored to one device, the same system works across many sensor types without retooling, automatically figuring out which sensor channel is cleanest and flagging bad data. This matters because it could let hospitals and at-home health devices reliably measure vital signs no matter what hardware is attached.
Technical view
CORAL reintroduces Short-Time Autocorrelation Functions (STACFs) to construct a 'correloform' — a 2D time-varying representation of signal periodicity — as a modality-agnostic backbone for rate estimation from quasi-periodic biosignals (ECG, PPG, SCG, BioZ). Because the transform is derived from generic autocorrelation mathematics rather than modality-specific filtering/feature engineering, the same pipeline yields rate estimation, noise resilience, automatic channel selection, and signal quality indicators across signal types without retraining. This is relevant to anyone building multimodal vital-sign monitoring pipelines (hospital or wearable) who wants a single robust front-end rather than maintaining per-modality DSP chains; benchmarking across in-hospital and at-home conditions suggests it generalizes across population and noise regimes.
arXiv · q-bio.PEConceptual
Watching a calm neighbor not flee is itself proof there's no danger — and groups exploit that math.
When an animal senses a possible predator, it faces a tradeoff: react fast and you'll often be wrong, react carefully and you might react too late. This paper shows that animals in a group can beat that tradeoff by paying attention not just to neighbors who bolt in fear, but also to neighbors who stay calm — because calm neighbors are quietly telling you 'I don't see a threat either.' The researchers modeled each animal as gradually gathering evidence of danger until it crosses a mental tipping point to flee, then compared a 'naive' animal that only reacts to others fleeing (which gets trigger-happy in bigger groups) against a 'smart' Bayesian animal that also credits calm as evidence of safety. The smart strategy lets the whole group escape real threats faster and with fewer false alarms, and it explains why panic can either fizzle out or cascade explosively through a herd depending on whether real danger is present.
Technical view
The authors model each individual as a drift-diffusion evidence-accumulator with a flee threshold, then formalize social inference by treating both neighbor flight (positive evidence) and neighbor stillness (negative evidence) as informative signals, collapsing the optimal Bayesian weighting into a single 'social discounting rate' parameter interpolating between naive (flight-only) and full Bayesian responders. This yields closed-form expressions for group detection speed, false-alarm rate, and cascade branching ratios, showing the branching ratio stays subcritical under safety but becomes supercritical under real threat — explaining self-limiting vs. explosive alarm cascades as an emergent property of the discounting rate rather than added mechanism. This provides a tractable analytical framework (rather than pure simulation) that collective-behavior and neuroscience-of-decision-making researchers could extend to empirical flocking/herding datasets or robotic swarm alarm protocols.
arXiv · cs.LGBuildable
A new algorithm finds shared patterns across datasets that don't even overlap directly, just via chained connections.
Scientists often have several different datasets — say gene activity, patient scans, and drug response data — that they want to analyze together to spot shared patterns, but standard methods require every dataset to share some common dimension, like the same patients or same genes measured in each. PathFinder gets around this by allowing datasets to connect indirectly: if dataset A shares something with dataset B, and B shares something with C, PathFinder can chain those overlaps into one combined analysis even though A and C never directly overlap. It works like a matchmaking network, tracing valid 'paths' through the web of shared connections to build one unified low-dimensional summary of everything. This matters because real-world data collected across different species, labs, or measurement types rarely lines up neatly, so this method could unlock joint analysis that used to be impossible.
Technical view
PathFinder extends joint/linked low-rank matrix decomposition methods (e.g., joint NMF/PCA variants) to settings where the full dataset collection lacks a common shared dimension, but pairwise or subgroup-level shared dimensions form a connected graph. By identifying paths through this graph of matrix-to-matrix shared dimensions, PathFinder propagates factor constraints transitively to estimate a global joint decomposition, recovering common latent patterns across modalities, species, or scales without requiring a universal one-to-one sample/feature mapping. This is directly applicable to multi-omics or cross-species integration pipelines where researchers currently must discard datasets that don't share an axis with every other dataset; implementation would build on existing joint-NMF/CCA-style optimization with graph-based path constraints.
arXiv · q-bio.QMRunnable
One Python toolkit finally lets you run and compare every RNA-folding algorithm without rewriting your code each time.
RNA molecules fold into shapes (secondary structures) that determine how they work in cells and how well drugs can target them, and scientists use many different computer algorithms to predict these shapes. The problem is that each algorithm has its own quirky input format, output style, and lack of built-in visuals, making it a hassle to compare methods or combine their results. MultiStructRNA solves this by giving researchers one simple interface that can run many different prediction algorithms, translate all their outputs into a shared, consistent format, and generate clear visualizations, whether working in a notebook or a large automated pipeline. This matters because it turns a fragmented, error-prone process into something researchers can trust and reproduce, especially for large-scale RNA studies used in drug and vaccine design.
Technical view
MultiStructRNA is a Python package that wraps multiple RNA secondary structure prediction algorithms behind a unified high-level API, standardizing heterogeneous inputs/outputs into a consistent object schema and computing ensemble-aware consensus metrics across predictors. It's built for both interactive (Jupyter) and production (pipeline/scripted) use, supports high-throughput batch analysis, and includes native visualization for structure comparison. Bioinformaticians can use it as a drop-in orchestration layer to benchmark or ensemble existing folding tools (e.g., comparing thermodynamic vs. comparative methods) without writing custom format-conversion glue code for each one.
arXiv · eess.ASRunnable
With just a handful of examples per call, a dead-simple 'nearest average' method can classify elephant sounds.
Researchers who want computers to recognize different types of elephant calls usually train complex classifiers on lots of labeled examples, but labeling animal sounds is slow and expensive. This study instead asks a simpler question: if you only have a few labeled examples per call type, how well does the most basic possible method work — just averaging the 'fingerprint' (embedding) of each known call type and matching new sounds to whichever average is closest? They test this simple approach using several pre-trained sound-recognition AI models as the fingerprint-makers, across real elephant vocalization datasets, and repeat the test many times with randomly resampled small example sets to check reliability. This matters because if a simple, parameter-free method works nearly as well as complex trained classifiers, it could make wildlife bioacoustics research faster and more accessible with far less labeled data.
Technical view
The authors evaluate nearest-centroid ('prototypical') classification on frozen pretrained acoustic embeddings (Perch v1, Perch v2, HuBERT-base layer 2, and MFCC baselines) for elephant call classification, using an N-way k-shot episodic protocol with class prototypes computed as the mean of support-set embeddings and query assignment via nearest centroid in squared Euclidean distance. Evaluated across the Elephant Voices and LDC datasets with 100-resample bootstrapping over support sets, this establishes a parameter-free lower-bound baseline against which trained classifiers can be benchmarked as labeled data scales from few-shot to full. Bioacoustics practitioners can use this as a near-zero-cost baseline pipeline (no training loop, just embedding extraction + centroid distance) before investing in fine-tuned classifiers for new species or call types.
arXiv · q-bio.PEConceptual
Random chance killing off rare species turns out to be what keeps whole ecosystems from collapsing.
Every species eventually goes extinct, but scientists have debated for decades whether having more species in an ecosystem makes it more or less stable overall. Most past models treated species survival as a fixed, deterministic outcome, ignoring the fact that small, rare populations can randomly die out just from bad luck (like a run of failed births), a phenomenon called demographic stochasticity. This paper builds a model that includes that randomness and finds it actually helps: random extinctions quickly weed out the weakest, lowest-population species, leaving behind a leaner ecosystem that's more stable overall. They also discover the pattern of how many species survive over time follows an unusual statistical shape with a 'heavy tail,' meaning occasional dramatic extinction events are more common than you'd expect. This reframes extinction not as pure loss but as a stabilizing pruning process in complex ecosystems.
Technical view
The authors incorporate demographic stochasticity into a rule-based large complex ecosystem model (extending classic random-matrix stability-diversity frameworks like May's), showing that stochastic extinction events preferentially prune low-abundance species and thereby increase the stability of the surviving community relative to deterministic-extinction baselines. They develop a bottom-up analytical theory characterizing extinction statistics and find the surviving-species fraction follows an anomalous heavy-tailed distribution rather than the exponential/Gaussian decay typically assumed. This offers theoretical ecologists a stochastic extension to classical stability-diversity theory and a testable statistical signature (heavy-tailed survival fraction) that could be checked against empirical community time-series or long-term ecological survey data.
arXiv · q-bio.GNConceptual
Ten practical rules to make gene-data science work without needing to see a single chart.
Modern biology research relies heavily on visual tools — colorful plots, heatmaps, and interactive charts — to make sense of huge datasets and decide what the results mean. But for blind and low-vision scientists using screen readers or braille displays, these visuals are often inaccessible, meaning the actual evidence behind a scientific decision is locked away in a format they can't use. This paper argues that fixing this accessibility gap has a bonus: it also makes research more reproducible for everyone, because both goals demand that every analytical decision be written down and explained in text rather than left implicit in a picture. The authors lay out ten concrete practices, like treating plots as documented decision records, using AI carefully to describe figures in words, and writing code and workflows that are text-first, so blind and sighted researchers alike can follow and verify the reasoning. This matters for making science genuinely open to more people and more trustworthy overall.
Technical view
The paper presents ten practical guidelines for non-visual bioinformatics aimed at researchers using screen readers, braille displays, or audio interfaces, framing visual analysis artifacts (QC plots, embeddings, heatmaps, genome browser tracks) as undocumented decision points that should instead be captured as explicit, text-based decision records. Recommendations span cautious use of AI-generated alt-text/figure descriptions, accessible computing environment setup, and text-first literate programming practices (e.g., structured markdown/notebook output over rendered-only graphics) that double as reproducibility documentation. Bioinformatics tool developers and lab leads can use this as a concrete checklist for auditing pipelines (e.g., ensuring every QC gate has a textual threshold/rationale, not just a plotted cutoff line) to improve both accessibility and audit-trail rigor.
arXiv · physics.soc-phConceptual
A model of "will you show up?" shows why shared plans can suddenly and permanently collapse.
This is a mathematical model of how people decide whether to do things together (like meeting a friend) or alone, when limited time is split between shared and private tasks. Each person weighs the value of the joint activity against the risk that the other might flake, so joining becomes a calculated bet rather than random chance. The researchers found that as this risk grows, the system doesn't drift gradually from "we do things together" to "everyone acts alone" — it snaps abruptly between the two, like a switch flipping. Once it falls into the solitary state it's very hard to recover from; someone has to stubbornly keep showing up alone for a while to pull the group back. This matters because it explains why social coordination in friendships, teams, or communities can break suddenly and stay broken.
Technical view
The authors extend the priority-queue model of human activity by letting agents endogenously set the priority of a joint task, discounted by their belief about a partner's participation probability, rather than drawing priorities from a fixed distribution. This strategic feedback introduces a phase transition between a "coupled" phase of sustained joint activity and an absorbing "solitary" phase, separated by a saddle-node bifurcation derived in closed form. The transition is discontinuous, and the solitary phase is absorbing so the system cannot spontaneously re-enter the coupled phase unless an agent unilaterally persists in the joint activity for roughly one memory-time, a cost the paper also quantifies. This gives a tractable analytical handle on hysteresis and irreversibility in coordination dynamics, applicable to modeling social-tie fragmentation or dyadic relationship breakdown.
arXiv · q-bio.PEConceptual
Bigger ants live longer, but size doesn't explain how fast they age or survive deadly heat.
Scientists wanted to know whether a worker ant's body size determines not just how long it lives, but also how it declines with age and how well it survives extreme heat — three questions usually lumped together as one. They studied 18 Australian ant species, tracking thousands of individual ants across paired field and lab settings to measure survival over time. Bigger ants reliably lived longer, likely due to basic physiology, while colony size and temperature had little effect on that pattern. Surprisingly, how fast an ant ages had nothing to do with size — instead it tracked when the species is active during the day, aging fastest in species active in the morning. This matters because it shows lifespan, aging, and heat tolerance follow different biological rules, which is important as climate change raises heat stress on ecosystems.
Technical view
Using paired field-laboratory survival assays across 18 Australian ant species (2,363 cohort-day observations, 1,148 workers), the study decomposes mortality risk into three components: lifespan duration, senescence trajectory, and thermal vulnerability. Cox proportional-hazards models show body size significantly predicts duration (HR = 0.67, p = 0.002) independent of colony size or a size×temperature interaction (both non-significant), with only a weak size×foraging-rate interaction (LRT p = 0.014) suggesting mostly intrinsic physiological drivers. Senescence trajectory instead tracks circadian niche (Kruskal-Wallis p = 0.009, steepest in matinal/morning-active species) rather than body size, decoupling aging rate from both lifespan and thermal death risk. The dataset offers a rare multi-species, multi-axis mortality decomposition useful for comparative life-history and climate-vulnerability modeling in social insects.
arXiv · eess.IVBuildable
A plug-in trick lets pathology AI skip boring tissue patches and focus only on the telling ones.
When AI analyzes a giant medical slide image, chopped into thousands of tiny patches, most patches are uninformative filler tissue that wastes computing power and can dilute the signal from the handful that actually matter for diagnosis. The researchers built TTIS, a method that at the moment of analysis automatically selects the small set of patches best representing the tissue, without any extra training — it just plugs into existing AI models. It also examines the slide from multiple views to make the selection more reliable. This matters because it could make cancer-diagnosis AI faster and sharper without the cost of retraining models on new hospital data.
Technical view
TTIS is a training-free, plug-and-play framework for whole slide image (WSI) multiple instance learning (MIL) that performs instance selection purely at inference time, choosing a compact, representative patch subset instead of processing every extracted patch. It adds a multi-view ensemble strategy that aggregates distinct facets of tissue morphology to improve selection robustness. Because it requires no retraining, it drops directly into any existing MIL pipeline (e.g., ABMIL, TransMIL, CLAM-style architectures) as an inference-time module, making it immediately deployable on pretrained pathology models. The practical payoff is reduced inference redundancy and compute alongside potential accuracy gains from filtering non-informative patches.
arXiv · eess.IVBuildable
Feeding pathology know-how into an AI's memory keeps it focused on the slide regions that matter.
This tackles the same giant-medical-image problem as before, but with Mamba, a newer AI architecture efficient at handling very long sequences of data like all the patches in a huge slide. The catch is that Mamba, looking only at raw visual features, can get distracted — irrelevant tissue crowds out the small but crucial diagnostic regions as it scans, diluting key evidence over the long "reading" process. The fix is to inject actual pathology knowledge into the AI's internal hidden memory state, nudging it to attend to what a pathologist would consider medically meaningful rather than just visually eye-catching. This matters because it could make AI slide analysis both faster, thanks to Mamba's efficiency, and more clinically trustworthy.
Technical view
KHiM-Mamba addresses a failure mode in selective state-space model (SSM/Mamba)-based MIL for WSIs, where purely vision-driven state updates misallocate attention across long patch sequences, causing the SSM's hidden state to accumulate task-irrelevant evidence and dilute diagnostically decisive regions. The proposed Knowledge-Aware Hidden-State Modulation architecture injects external pathology knowledge directly into the Mamba hidden state, biasing its selective update/readout dynamics toward clinically meaningful regions rather than relying solely on visual saliency. This preserves Mamba's core advantage — linear complexity over long sequences — while correcting its knowledge-blind selectivity, positioning it as an alternative to attention-based MIL aggregators for gigapixel WSI encoding. Practitioners on SSM-based pathology pipelines could adopt the hidden-state modulation mechanism as a general way to fuse domain priors into SSM state dynamics beyond WSI analysis.
arXiv · q-bio.PEConceptual
A clever simplification finally lets scientists write exact formulas for why epidemics wax and wane in cycles.
Diseases with multiple strains, like flu variants, often arrive in repeating waves rather than settling into steady state, but the math describing two interacting strains has been too messy to solve exactly. The researchers found a special, simplified version of the two-strain model — where immunity to one strain fully blocks second infections from it in an asymmetric way — that turns out to be exactly solvable with clean formulas. Using this, they precisely map when the disease settles into steady coexistence versus when it breaks into self-sustaining oscillating outbreaks, driven by how much more infectious "second-time" infections are. Simulations further reveal that small everyday fluctuations and big dramatic outbreak cycles can occur side by side in the same system. This matters because it gives epidemiologists exact mathematical tools, not just simulations, for understanding multi-strain disease cycles like flu or dengue.
Technical view
The paper identifies an analytically tractable subclass of two-strain SIR-type models with asymmetric cross-immunity, where immunity to one strain fully blocks secondary infections from that strain, eliminating the algebraic complexity that has historically blocked closed-form analysis of multi-strain coexistence. Within this class, the authors derive explicit expressions for coexistence equilibria and their stability, showing the relative transmission advantage of secondary infections is the key bifurcation parameter governing onset of limit-cycle oscillations, generalizing prior narrow near-threshold results to a broader regime. Numerical simulation of the full (non-simplified) system reveals a richer bifurcation landscape than the tractable subcase alone predicts, with small-amplitude local oscillations coexisting with large-amplitude recurrent outbreak cycles. This offers epidemiological modelers a reference analytical benchmark for validating numerical multi-strain models and for reasoning about strain-competition-driven oscillatory dynamics such as influenza subtype cycling or dengue serotype dynamics.
arXiv · q-bio.NCConceptual
A survey maps the toolbox of data methods racing to catch brain disease before it's too late.
Diseases like Alzheimer's and Parkinson's are usually diagnosed only after a lot of irreversible brain damage has occurred, so there's a big push to spot subtle warning signs earlier using brain scans and data analysis. This paper is a review that rounds up a wide range of modern data-driven methods, organized into four main categories, that researchers use to detect small, early, person-specific brain changes. Rather than inventing a new method, it maps the current landscape, showing how these diverse approaches all aim at the same goal: personalized models of brain health grounded in real biology and useful in the clinic. It also lays out what's still unsolved statistically, computationally, and clinically. This matters because it could help clinicians catch neurodegenerative disease earlier, when treatment might still make a difference.
Technical view
This review surveys data-driven methods for translational neuroscience and personalized neuro-health, organized around four methodological pillars spanning neuroimaging-based quantitative biomarker detection for early, individualized neurodegenerative change. Its throughline is convergence: despite methodological diversity, these approaches share the translational goal of producing personalized, mechanistically grounded, clinically actionable brain-health models for diseases like Alzheimer's and Parkinson's, currently diagnosed only after substantial irreversible neuronal loss. It closes by cataloging open statistical, computational, and clinical challenges, making it a useful orientation point for researchers deciding which methodological pillar — specific imaging biomarkers, modeling frameworks, etc. — to build on for early-detection or precision-neurology work. As a review, its value lies in synthesis rather than a novel result to replicate directly.
arXiv · q-bio.PEConceptual
A math model shows why cells sometimes deliberately split their emergency supplies unevenly between offspring.
When a cell divides, it must decide how to split a limited protective reserve — like a rainy-day fund against stress — between its two daughter cells. You'd assume splitting evenly is always safest, but this model shows that when environmental risk is shared between the daughters (say, they face the same conditions), giving unequal amounts to each can actually be better for long-term survival. Using a mathematical measure of long-run growth rate under randomness, the researchers find this unequal-split strategy suddenly becomes favored once environmental sharing crosses a threshold, happening as an abrupt jump to a specific asymmetry level rather than a gradual drift. This holds up even when other model details are changed. This matters for understanding why asymmetric division shows up in biology, such as in stem cells or bacteria.
Technical view
The paper models cell division as reserve-partitioning between mother and two daughters, where each fixed partition policy (parameterized by asymmetry α) generates a random demographic operator whose top Lyapunov exponent determines the population's long-term growth rate under environmental fluctuations. Environmental sharing between sibling lineages is coupled to the value of diversification: weak sharing favors symmetric (α≈0.5) inheritance, while sufficiently strong shared risk selects a distinct asymmetric branch near α≈0.2, with the transition occurring as a discontinuous jump between fitness-optimal branches rather than a continuous bifurcation. This asymmetric-optimal phase is robust to changes in the protection law and reserve turnover dynamics, though the exact transition boundary depends on protection nonlinearity and reserve memory — a Lyapunov-exponent framework that could be extended to test specific molecular reserve systems, such as damage or chaperone partitioning, against measured asymmetry ratios.
arXiv · cs.CLBuildable
Teaching AI to zoom in on the right part of a molecule before reasoning speeds it up nearly 10x.
To predict a molecule's properties or suggest edits, an AI needs to know which specific parts of its structure matter chemically — but current models either get told which parts matter by humans or must guess properties from a full image without knowing where to look. VLSR teaches the AI to first spot the chemically important regions in a picture of the molecule on its own, then reason about how those regions affect its properties inside a compact internal workspace, rather than describing everything at length. It's like teaching someone to circle the important part of a diagram before explaining what it does, instead of narrating the whole picture. This localize-then-reason approach lets the AI process molecules 9.6 times faster than a comparable baseline. This matters for drug discovery and materials science, where AI needs to reason about chemical structures quickly and accurately at scale.
Technical view
VLSR (Visual Latent Structural Reasoning) is an end-to-end framework for molecular property and edit reasoning that jointly learns localization of chemically meaningful regions in a molecular image and property reasoning over those regions, rather than relying on externally provided motif annotations or reasoning directly and unstructured from full images. Its core "localize-then-reason" strategy first learns region localization, then performs property-effect reasoning in a compact latent workspace before decoding the final answer, avoiding verbose token-level chain-of-thought over raw pixels. Under matched inference settings, this design achieves 9.6x higher throughput than a comparison LLM-based baseline, suggesting the latent reasoning stage substantially cuts inference cost versus token-heavy chain-of-thought approaches. This is directly relevant to practitioners building faster multimodal chemical-reasoning or molecular-editing pipelines who need throughput gains without sacrificing localization-grounded interpretability.
arXiv · physics.ins-detBuildable
Electron microscopes blink faster than their steering coils can keep up, blurring the picture.
In advanced electron microscopy called 4D-STEM, a beam of electrons is scanned across a sample point by point, snapping a tiny diffraction pattern at each spot. As scientists push to faster scan speeds (microseconds per point) to reduce damage to delicate samples like proteins, they discovered the electromagnetic coils that steer the beam can't quite keep up in time, causing the beam to lag and smear the image in one direction. The fix is a clever computational trick: compare shifted sub-frames of the data to measure exactly how much lag occurred, then mathematically undo the smear. This matters because it lets scientists recover crisp, trustworthy images from fast, low-dose scans without buying new hardware, which is especially valuable for imaging fragile biological samples.
Technical view
The authors identify an intra-dwell scan-coil settling delay (tens of microseconds) that becomes non-negligible at microsecond dwell times used in fast pixelated-detector 4D-STEM, producing anisotropic signal smearing along the fast-scan axis. They characterize this via direct probe imaging and sub-frame diffraction analysis, then correct it using phase-correlation-based sub-frame alignment applied post-hoc to existing datasets. The correction recovers signal across spatial frequencies, with the largest benefit at large scan steps typical of low-dose biological 4D-STEM. Because it requires no hardware modification, practitioners can apply this correction retroactively to archived fast-scan datasets to improve reconstruction fidelity.
arXiv · q-bio.GNBuildable
AI rewrote a scientist's clunky old bioinformatics code into Rust — 80x smaller, way faster.
Bioinformatics — the software used to analyze DNA and medical images — is often built on decades-old code in languages like Perl or Fortran that few people still know how to maintain, creating security risks and wasted computing resources. The researchers used an 'agentic' AI system (one that can plan and act autonomously, guided by automated code-checking tools) to translate this legacy code into Rust, a modern, fast, and safety-focused programming language. Applied to their own tool, Bascet, the translated version ended up about 80 times smaller, built ten times faster, and ran key operations three times quicker, while also shedding messy external dependencies. This matters because it offers a practical path to modernize critical scientific software without a costly, error-prone manual rewrite.
Technical view
The authors combine static analysis with agentic AI (an LLM-driven system capable of iterative planning and tool use) to systematically translate legacy bioinformatics codebases into Rust, releasing prompts and supporting software for the pipeline. Applied to their NGS/imaging tool Bascet, the translation achieved ~80x reduction in codebase/binary size, ~10x faster build times, and >3x runtime improvement on key steps, while eliminating Unix-specific dependencies for better portability. This demonstrates a reproducible, static-analysis-guided methodology practitioners could apply to other legacy scientific codebases (e.g., Perl, Fortran) facing similar maintainability and safety concerns. The approach suggests a template for using AI agents as verified code-migration tools rather than purely generative assistants.
arXiv · q-bio.PEConceptual
A simulated bird-flu outbreak on a fake island tests when it's safe to restock chicken farms.
When bird flu hits poultry farms, officials must decide fast how aggressively to cull animals and how soon depopulated farms can safely restart operations. This study builds a computer simulation of a fictional 'Jolly Island' with different farm types (broiler chickens, organic ducks, and others), modeling how the virus spreads through nearby farms, the environment, animal transport, and distance, then testing control strategies like preventive culling and delayed restocking. The simulation showed that culling farms preemptively — before they're confirmed infected — cut the total outbreak burden by about 17%, and confining birds earlier shrank the epidemic further. This matters because it gives policymakers an evidence-based way to weigh the economic and animal-welfare costs of aggressive intervention against the risk of letting an outbreak spread.
Technical view
The authors construct a stochastic spatial SEIR (susceptible-exposed-infectious-recovered) metapopulation model for a synthetic HPAI outbreak, incorporating local, environmental, movement-mediated, and distance-dependent transmission across farm types (Broiler-2, organic duck, Other), alongside reactive/preventive culling, production-specific confinement, and capacity-based restocking rules. Simulations show geographically concentrated epidemics that vary substantially by production class, with preventive culling reducing mean cumulative burden from 16,362.7 to 13,631.9 infectious-farm-days (16.7% reduction), and earlier confinement further reducing epidemic magnitude. The modular framework (synthetic geography, class-specific transmission parameters, intervention triggers) is designed to be adapted to real regional poultry networks for scenario testing. Practitioners in animal health policy could use this structure to benchmark culling thresholds and restocking capacity constraints before applying to real outbreak data.
arXiv · q-bio.NCConceptual
Alzheimer's-causing proteins seem to hitch a ride on brain activity itself as they spread.
Alzheimer's and similar diseases progress as toxic misfolded proteins spread from one brain region to connected ones, almost like an infection traveling along neural highways. Prior math models of this spread ignored the fact that neurons firing electrical signals actually helps push these proteins along, something lab experiments have shown. This paper builds a model that links neuronal activity to protein-spreading dynamics (borrowed from epidemic math used to study diseases like flu), finding a tipping point that determines whether a tiny pathological seed grows into full-blown spread, and how activity redirects which brain areas get hit first. This matters because it could help predict individualized disease progression and identify why certain brain networks are more vulnerable, potentially informing where and how to intervene.
Technical view
The authors couple a generic node-activity process to susceptible-infected-susceptible (SIS) epidemic dynamics on brain connectomes, deriving an epidemic threshold that governs whether pathological protein seeds propagate, and showing that a dominant network eigenmode determines the spatial origin of growth. Analytical approximations quantify how neuronal activity shifts this threshold and reshapes spreading trajectories by mixing structural network modes (eigenvectors of the connectivity matrix). For networks with multiscale (hierarchical/modular) structure, they decompose activity-driven effects into contributions from regional mean activity versus within-region variability, enabling attribution of spreading changes to specific network scales. This provides a mechanistic, mode-decomposition framework that researchers could apply to patient-specific connectomes and activity data (e.g., fMRI, EEG) to predict individualized Alzheimer's progression patterns.
arXiv · q-bio.PEConceptual
Bacteria don't just face feast or famine — they navigate many shades of 'okay, not great.'
Microbes living in changing environments — think gut bacteria or lab cultures — often face swings between plenty of food and scarcity, and scientists have long modeled this as a simple on/off 'feast or famine' switch. But real environments are messier, with many in-between levels of resource availability, not just two extremes. This paper studies how two competing bacterial strains, one growing slightly slower than the other, fare when the environment cycles through several intermediate states rather than just two, each state offering a different capacity for how many microbes it can support. This matters because understanding these richer fluctuation patterns could reveal why weaker competitors sometimes survive or even thrive in fluctuating real-world conditions, with implications for microbiome ecology and evolution.
Technical view
The authors extend classic binary switching-environment models (used to study feast-famine population dynamics) to a multi-state stochastic framework where the environment transitions among a finite number of intermediate states, each with its own carrying capacity, better approximating experimentally observed gradual nutrient fluctuations. They analyze competitive dynamics between two strains with slightly different growth rates under this richer switching process, likely deriving conditions (switching rates, state structure) under which the slower strain persists or is driven extinct. This generalizes prior two-state (Moran-type or telegraph-process) population models to a broader class of environmental noise, providing a more realistic null model for microbial competition. Researchers modeling real fluctuating ecosystems (e.g., gut microbiota, chemostats) could apply this framework to test how granularity of environmental variation affects coexistence outcomes.
arXiv · cs.LGBuildable
AI models of proteins are usually read from their 'last thought' — but earlier layers know more.
Protein language models are AI systems trained on huge databases of protein sequences, similar to how ChatGPT is trained on text, and they convert amino acid sequences into number-based representations used for tasks like predicting protein structure or function. Everyone conventionally uses the output from the model's very last processing layer, treating it like the model's 'final answer,' but nobody had carefully checked whether that's actually the best layer to use. The researchers tested 13 different protein language models across 15 different tasks, training simple predictor 'probes' on the outputs of each internal layer, and found that the last layer is rarely the most informative one. This matters because it means researchers using these models for drug discovery or protein engineering may be leaving performance on the table by defaulting to the last layer instead of the best one.
Technical view
The authors systematically probe 13 protein language models (PLMs) across 15 downstream tasks drawn from 11 datasets, training linear/shallow probes on embeddings extracted from every intermediate layer rather than only the conventional final layer, and supplement this with latent-space geometry characterizations to estimate embedded information content. The key finding: final-layer embeddings rarely yield the best downstream task performance, challenging the field's default convention. This suggests a practical, low-cost intervention — layer selection via lightweight probing — that practitioners can apply to existing pretrained PLMs to improve downstream performance without retraining, and motivates further mechanistic interpretability work on what different PLM layers encode.
arXiv · stat.MEConceptual
A clever statistical trick tests whether brain-wave predictions are cheating by peeking at the future.
When scientists build a model that predicts a brain signal (like an EEG wave that ramps up in anticipation of an event), it's hard to know if the model is genuinely using only past information, or if it's secretly benefiting from information that technically comes later. This paper proposes a rigorous testing method: lock in the prediction endpoint before randomly varying the delay to the actual event, so that if the model truly only used past-available information, its predictions shouldn't be able to reliably track that later-randomized delay. Applying this to a specific anticipatory brain wave (called contingent negative variation), they build in strict safeguards against accidental data leakage to make the test trustworthy. This matters because it offers neuroscientists (and other predictive modelers) a principled way to catch models that look good on paper but are secretly relying on information they shouldn't have access to.
Technical view
The paper introduces 'Level II-A,' a design-based causal inference framework that distinguishes model fit from information sufficiency by using post-endpoint randomization: a pre-event endpoint is committed before the delay-to-imperative-event is randomized, turning the later-assigned delay into a negative-control probe. Under a 'past-adapted factorization,' any model using only pre-commitment information should be unable to systematically order the endpoint by the randomized delay; violation of this indicates the model leveraged information beyond what was legitimately available. Applied to anticipatory EEG (contingent negative variation), the framework combines leakage-safe preprocessing with a frozen, label-blind comparator and retained-sample qualifications to produce a confirmatory residual test. Researchers building predictive models from time-series neural data can adopt this design to rigorously test sufficiency claims rather than relying on fit-based validation alone, which is vulnerable to inadvertent information leakage.
arXiv · q-bio.NCConceptual
Could a math formula ever let you truly know what it's like to be someone else's mind?
Philosophers have long wrestled with the 'problem of other minds': you know your own inner experience directly, but you can only ever infer someone else's from the outside, leaving what the authors call an 'acquaintance gap.' Some researchers hope that by describing conscious experience mathematically — essentially finding a universal translator or 'Rosetta Stone' for subjective experience — we could bridge that gap. This chapter examines two competing mathematical frameworks for consciousness, the Qualia Structure Paradigm and Integrated Information Theory, asking what each would actually let us conclude about another being's inner experience if it succeeded on its own terms. This matters for debates about animal consciousness, AI sentience, and medical assessments of unresponsive patients, where knowing whether 'someone is home' has real ethical stakes.
Technical view
The chapter conducts a comparative philosophical analysis of two structuralist theories of consciousness — the Qualia Structure Paradigm (Qstr) and Integrated Information Theory (IIT) — evaluating their capacity to license principled inference about another system's phenomenal experience via a mathematical 'Rosetta Stone' translating experiential content into formal structure. Qstr is characterized as proceeding inter-phenomenally, aiming to exhaustively characterize experience through its internal relational structure, while IIT is presumably contrasted on its integrated-information-theoretic mechanism (the abstract cuts off before full elaboration). The analysis likely interrogates whether structural isomorphism between a formal model and a system's causal/relational architecture is sufficient to warrant inference to genuine phenomenal states, bearing on consciousness-detection frameworks proposed for AI systems, animals, or disorders of consciousness. Readers building or evaluating consciousness-measurement frameworks (e.g., IIT's Φ metric) would find this a rigorous epistemic critique of what such measures can and cannot license inferentially.
bioRxiv · molecular biologyConceptual
A misfolded liver protein doesn't just clump — it quietly wrecks cells' power plants and fat processing.
Alpha-1 antitrypsin deficiency is a genetic disease where a faulty version of a liver protein, called Z-AAT, misfolds and gets stuck inside liver cells instead of being released into the blood. Researchers grew both flat liver cell cultures and tiny 3D lab-grown mini-livers (organoids) from patients' own cells to watch what the stuck protein does over time. They found it doesn't just pile up — it also causes fat to accumulate, damages mitochondria (the parts of the cell that make energy), and forces cells to rely more on sugar than fat for fuel. This matters because it reframes the disease as a broader metabolic breakdown, not just a protein traffic jam, opening new angles for treatment.
Technical view
Using Z-HepG2 cells and ZZ patient-derived hepatic organoids, the authors combined transcriptomic and proteomic profiling with functional assays of mitochondrial and peroxisomal dynamics to characterize downstream effects of Z-AAT polymer accumulation. Z-AAT expression reduced protein secretion and drove lipid accumulation alongside mitochondrial structural abnormalities, an increase in mitochondrial number, and impaired respiratory capacity, with metabolic profiling showing reduced oxidative phosphorylation and a compensatory shift toward glycolysis. The organoid model provides a patient-relevant platform for testing whether restoring mitochondrial or lipid handling mitigates AATD-associated liver disease. This positions organelle-level metabolic dysfunction, not just proteotoxic stress, as a therapeutic target.
bioRxiv · cell biologyConceptual
Losing one copy of a genome-guarding protein leaves specific chromosome spots prone to snapping.
Cells have a family of three related 'HP1' proteins that help pack DNA safely and keep chromosomes stable during division. This study knocked out each HP1 type one at a time in cells and then stressed their DNA-copying machinery with a chemical, checking under a microscope for broken chromosomes. They discovered that losing just one specific version, HP1-alpha, causes breaks at particular genome locations, while losing a very similar cousin, HP1-beta, does not — showing these lookalike proteins actually have distinct, non-swappable jobs. The mechanism seems to be that HP1-alpha loss slows down the process of copying DNA, making certain fragile regions more likely to snap under stress, which is relevant to understanding cancer-related chromosome instability.
Technical view
The authors performed isotype-specific inactivation of HP1α versus HP1β across multiple cell lines and quantified chromosomal breakage on metaphase spreads with and without aphidicolin-induced replication stress. HP1α loss, but not HP1β loss, significantly increased breaks on chromosome arms and within pericentromeric heterochromatin, and was mechanistically linked to reduced replication fork velocity. This defines a set of genomic loci that behave as HP1α-dependent common fragile sites, distinct from classical fragile sites, giving researchers a new isotype-specific readout for probing heterochromatin's role in replication stress and genome instability.
bioRxiv · cell biologyBuildable
A new imaging trick zooms into single molecules inside a parasite without losing the big picture.
Scientists studying cells under powerful microscopes face a tradeoff: some methods show incredible molecular detail but only in a tiny patch of tissue, while others capture a whole cell but blurrily. This project built a new workflow, called VolWeaver, that lets researchers pick a specific tiny region inside a malaria parasite and image it at near-atomic sharpness while still keeping several micrometers of surrounding cellular context visible. It works by carefully optimizing a technique called transmission electron tomography, which takes many angled snapshots of a sample and reconstructs them into a 3D volume. This matters because it finally lets scientists connect fine molecular structures to the broader architecture of the cell they live in, which is crucial for understanding how parasites like the one causing malaria actually work.
Technical view
VolWeaver is an optimized transmission electron tomography (TEM) pipeline for resin-embedded, large-scale volume imaging that targets specific regions of interest at nanometer resolution while retaining several micrometers of surrounding cellular context, applied here to malaria parasites. It bridges the resolution/field-of-view gap between single-particle cryo-EM/cryo-ET and lower-resolution volume EM (e.g., resin-embedded SEM), enabling correlated multiscale analysis within one specimen. Practitioners working with resin-embedded pathogens or organelle-scale structures could adopt the workflow to link molecular-scale tomographic detail directly to whole-cell ultrastructure without switching modalities or samples.
bioRxiv · developmental biologyConceptual
Fruit-fly testis cells use a directional actin 'push' to build the nursery where stem cells live.
Stem cells throughout the body depend on a specialized support structure called a niche, but scientists rarely get to watch a niche actually form because it's usually deep inside developing tissue. Using the fruit fly testis as a visible model system, researchers tracked a cell scaffolding protein called F-actin, which becomes lopsidedly concentrated at specific cell-to-cell contact points as the niche assembles. To test whether this lopsided pattern actively drives cell movement or is just a byproduct of cells sticking together, they used a light-controlled ('optogenetic') tool to switch off a regulator called Rho1 at precise moments and locations, disrupting the actin pattern on demand. Doing so broke normal niche formation, showing that the directional actin arrangement is a cause, not just a consequence, of how stem cell niches get built.
Technical view
Using live imaging of Drosophila testis niche formation, the authors show that F-actin polarizes to specific cell-cell interfaces during niche assembly, and they test causality using optogenetic disruption of Rho1 to acutely and locally perturb cortical F-actin with tissue and temporal precision. Rho1-mediated loss of F-actin polarization produced defects in niche formation, indicating that polarized actin actively directs niche cell positioning/motility rather than merely reflecting adhesion-driven sorting. This establishes an in vivo, optogenetically tractable system for dissecting the cytoskeletal mechanics of stem cell niche morphogenesis, applicable to studying niche construction principles more broadly.
bioRxiv · biochemistryConceptual
An enzyme crystal's own water content secretly dictates how much the protein inside can wiggle.
X-ray crystallography lets scientists see the 3D shape of proteins, but it usually freezes them into one rigid pose, when in reality proteins constantly flex and move. This study used an extremely fast, high-powered X-ray laser technique to look at many tiny crystals of a plant enzyme called lipoxygenase, hoping to catch its natural range of motion. While analyzing the data, they noticed the crystals weren't all identical — some had more water packed between the protein molecules than others — and this water content directly changed how flexible the protein appeared to be. By carefully sorting the data into two groups based on this hidden variation, they solved two distinct structures from a single experiment, revealing that the crystal's surroundings, not just the protein itself, shape what scientists see.
Technical view
Using serial femtosecond crystallography (SFX) at the LCLS free-electron laser on soybean lipoxygenase-1 microcrystal slurries, the authors identified unit-cell polymorphism arising from both indexing ambiguity (pseudo-tetragonal lattice symmetry) and genuine non-isomorphism tied to differing crystal solvent content. By combining unit-cell clustering with systematic reindexing, they deconvolved the mixed dataset into two polymorphs and solved two independent structures showing distinct conformational states from a single experiment. This provides a general data-processing strategy—clustering plus reindexing—for extracting hidden conformational heterogeneity from SFX datasets that would otherwise be merged and averaged away, relevant to anyone doing room-temperature/physiological serial crystallography.
bioRxiv · biochemistryConceptual
Where blood swirls roughly in an artery is exactly where its proteins start looking diseased.
Fatty plaques that clog arteries don't form randomly — they tend to show up at spots where blood flow gets turbulent or disturbed rather than flowing smoothly. Researchers wanted to know what's different, at the protein level, between plaque-prone and plaque-resistant spots in the same blood vessel, but this required measuring proteins in extremely small tissue samples, which used to be technically impossible. Using improved, highly sensitive lab equipment (mass spectrometry) that can now analyze tiny amounts of tissue, they compared proteins from plaque-forming branch points versus healthy-looking curves of the aorta in mice bred to develop atherosclerosis. This location-by-location protein mapping reveals a gradient of disease-related changes tied directly to how blood physically moves through the vessel, which could point to new early-warning markers or treatment targets.
Technical view
The study performed site-resolved proteomics on aortic arch tissue from Western-diet-fed ApoE-/- mice, comparing plaque-prone regions at major branch points and the inner curvature (disturbed flow) against visibly healthy, flow-protected regions, using LC-MS/MS enabled by recent advances in low-input mass spectrometry. This generates a spatial proteomic map linking local hemodynamic conditions to site-specific protein signatures of atherosclerotic disease progression within a single animal. Researchers can use this dataset to identify candidate mechanotransduction pathways or biomarkers specific to flow-driven plaque initiation, and the low-input MS approach itself is reusable for other small-tissue-sample proteomic studies.
bioRxiv · bioengineeringBuildable
A squeezed-light camera trick films a beating heart in 3D, 600 times a second, through living tissue.
Seeing fast 3D movement deep inside living tissue is hard because normal volumetric microscopes either have to scan slowly point-by-point, or they split their camera's limited pixels across many viewpoints at once, losing detail. This is especially tough in a useful infrared wavelength band (NIR-II) that penetrates tissue well but whose cameras are small and noisy. The researchers built a new microscope, NIR-II squeezed light-field microscopy, that optically twists and compresses multiple viewing angles together before they hit the camera, so the limited pixels get used far more efficiently while still capturing enough information to reconstruct a 3D image. The result is volumetric video fast enough — up to 600 3D frames per second — to capture things like a beating heart in real time, without needing any fluorescent labels or dyes.
Technical view
NIR-II SLIM introduces an optical rotate-and-compress step for multi-perspective light-field views prior to detection, maximizing effective use of a small-format, high-noise InGaAs sensor in the second near-infrared window while preserving the spatial information needed for 3D reconstruction. The system achieves volumetric acquisition rates up to 600 volumes/s at a 512×512 reconstructed lateral sampling grid, demonstrated on label-free, four-dimensional imaging of cardiac dynamics in vivo. This addresses the core pixel-budget bottleneck of snapshot light-field imaging in NIR-II, offering a hardware-optical (rather than purely computational) route to high-speed deep-tissue volumetric imaging that could be adapted to other fast physiological dynamics.
bioRxiv · bioinformaticsRunnable
An algorithm spots tissue cells that broke their normal honeycomb formation — a red flag for disease.
Healthy tissues like skin or gut linings have cells arranged in tidy, repeating hexagonal patterns, and when disease starts, that neat organization breaks down into oddly placed cells with abnormal gene activity. The researchers built a computer tool called SPADE that combines two types of data — regular single-cell gene readouts and spatial maps showing where cells actually sit in tissue — to learn what 'normal' spatial-and-genetic patterns look like. It then flags locations that don't fit that learned normal pattern, while also reporting how confident it is in each flag, so users know which anomalies to trust versus which might be noise. This gives researchers a statistically principled way to automatically catch early signs of disease-related tissue disorganization instead of relying on visual inspection.
Technical view
SPADE integrates scRNA-seq and spatial transcriptomics via a variational autoencoder combined with Gaussian mixture modeling to jointly learn cell-type embeddings and perform spatial deconvolution, then applies conformal prediction to yield uncertainty-calibrated detection of spatially aberrant spots. The conformal prediction layer provides formal statistical coverage guarantees on aberrancy calls rather than ad hoc thresholding, and the method reportedly outperforms existing approaches in benchmark validation. As a computational framework it's directly usable by anyone with paired scRNA-seq/spatial transcriptomics datasets to screen for disease-associated spatial disorganization, likely released as an installable analysis pipeline.
bioRxiv · cancer biologyConceptual
Leukemia cells' DNA-copying machinery can be sabotaged to make cancer drugs work better.
Acute myeloid leukemia (AML) is caused by bone marrow cells multiplying out of control, and current treatments are often harsh and don't always work. This research looks at drugs called PARP inhibitors, which block a DNA-repair protein, and finds they only work well in some genetic subtypes of AML. Using detailed imaging of individual DNA strands as cells copy their genome, the researchers show that in drug-sensitive leukemia cells, the treatment first speeds up DNA copying dangerously fast, then causes the DNA strand to snap later in the same cell cycle. Understanding this timeline could help doctors predict which patients will actually benefit from these drugs.
Technical view
The study combines single-cell and single-molecule DNA fiber assays with damage-signaling readouts to dissect replication fork dynamics under PARP inhibition and cytarabine in AML models stratified by fusion genotype. In PARPi-sensitive backgrounds (RUNX1-RUNX1T1, PML-RARα), PARP inhibition deregulates RECQ1-dependent fork restart, producing an initial phase of fork acceleration followed by fork breakage within the same S phase, whereas KMT2A-rearranged (resistant) cells apparently evade this cascade. This mechanistic ordering — accelerated restart preceding catastrophic breakage — offers a candidate biomarker/timing window for combination scheduling and could guide biomarker-driven PARPi trial stratification in AML.
bioRxiv · cancer biologyConceptual
One protein makes breast cancer cells drink in more chemo instead of pumping it back out.
Breast cancer chemotherapy doesn't work equally well for everyone, partly because some tumor cells act like resistant 'stem cells' that survive treatment and cause relapse. This study looks at a protein called p140Cap and finds that it helps chemotherapy drugs like doxorubicin stay inside cancer cells longer, causing more DNA damage and killing more cells. It works by blocking a signaling pathway (Wnt/β-Catenin) that would otherwise boost a pump protein (ABCC1) that stem-like cells use to flush the drug back out. This suggests that boosting p140Cap activity, or blocking the Wnt pathway it controls, could make existing chemotherapies more effective against hard-to-treat breast cancers.
Technical view
Using preclinical and patient-derived HER2-positive and triple-negative breast cancer models, the authors show p140Cap increases intracellular doxorubicin retention and downstream DNA damage/apoptosis by restraining a doxorubicin-negative, ABCC1-high side population enriched for stem-like cells. Mechanistically, p140Cap inhibits β-Catenin signaling, which otherwise drives ABCC1 (a drug-efflux transporter) expression; constitutively active β-Catenin reverses the chemosensitizing phenotype, while pharmacological Wnt/β-Catenin inhibition phenocopies p140Cap's effect. This positions the p140Cap–β-Catenin–ABCC1 axis as a druggable node for combination strategies aimed at depleting the chemoresistant stem-like compartment.
bioRxiv · plant biologyConceptual
Knock out one type of plant water channel, and a related one quietly gets destroyed too.
Plant cells regulate water flow using channel proteins called aquaporins, which come in two related families, PIP1 and PIP2, sitting in the cell membrane. Researchers genetically removed several PIP2 versions in the mustard-family plant Arabidopsis and unexpectedly found that PIP1 protein levels also crashed, even though the genetic instructions (mRNA) for making PIP1 were untouched. This means the loss happens after the protein is made, not because the plant stopped producing it, hinting that PIP1 proteins need PIP2 partners to avoid being tagged for destruction by the cell's protein-disposal systems. It's a reminder that in biology, removing one gene can ripple through and silently take out a partner protein via a completely different mechanism.
Technical view
In a pip2;1 pip2;2 pip2;4 pip2;6 pip2;7 quintuple mutant, Arabidopsis PIP1 protein abundance is strongly reduced despite unchanged PIP1 steady-state transcript levels and polysome loading, indicating a post-translational (not transcriptional/translational) mechanism. Intermediate mutant combinations (pip2;1 pip2;2 and pip2;1 pip2;2 pip2;7) show graded residual PIP1 protein (60% and 20%), consistent with dose-dependent stabilization of PIP1 by PIP2 heteromerization. The authors are positioned to test which degradation pathway — ER-associated ubiquitin-proteasome degradation or an alternative route — clears unpartnered PIP1, informing models of aquaporin heteromer-dependent quality control at the plasma membrane.
bioRxiv · plant biologyBuildable
Cow manure plus a bit less irrigation water can grow the same wheat with less nitrate pollution.
Farmers in dry regions face a tricky balance: use enough fertilizer and water to grow good crops, without wasting water or letting nitrogen runoff pollute groundwater. This two-year field study in Pakistan compared spreading dairy manure versus synthetic urea fertilizer, combined with either full or reduced ('deficit') irrigation, across a wheat-fallow-maize rotation. They buried special sensors deep in the soil to directly measure how much nitrate leaked downward past the root zone, and used a computer model to estimate daily water loss. The combination of manure with reduced irrigation turned out to be a promising sweet spot, suggesting practical changes farmers could make to protect water quality without sacrificing yield.
Technical view
The two-year field trial crossed dairy manure (50 Mg/ha annually, N-equivalent to recommended urea) against sole urea fertilization with 100% vs 75% ETc irrigation in a wheat-fallow-maize rotation, using suction lysimeters at 1.2 m depth for direct nitrate-N leachate measurement and HYDRUS-1D modeling for daily deep percolation. A significant manure × irrigation interaction affected yield, nitrate-N leaching, and soil quality, with manure under deficit irrigation improving wheat outcomes — data practitioners could use to parameterize regional HYDRUS-1D models or design manure/deficit-irrigation protocols for semi-arid cereal systems.
bioRxiv · plant biologyRunnable
'Forever chemicals' in water can stunt bean sprouts before they even get their first leaf.
PFAS, nicknamed 'forever chemicals' because they don't break down in the environment, are showing up in soil and water everywhere, but scientists don't fully know how they affect the crops we eat. This study grew mung beans in water spiked with two common PFAS chemicals (PFOA and PFOS) at different concentrations, to see how the plants responded. At very high doses, PFOA badly delayed seed sprouting, leaf growth, and root hair formation, while a wider range of doses caused temporary stunted growth that plants partly recovered from over time. At the highest tested dose, plants ended up smaller and lighter, showing that PFAS contamination can measurably harm crop development, and worse at higher concentrations.
Technical view
Hydroponic mung bean (Vigna radiata) seedlings were exposed to PFOA and PFOS across a dose range (5-500 µM, plus a 1 mM PFOA high-dose condition), with developmental endpoints including germination timing, leaf emergence, root hair formation, wet weight, and leaf area/biomass tracked over multiple timepoints (48 h onward). PFOA at 1 mM produced severe, more pronounced impairment than PFOS at equivalent exposure, while the 5-500 µM range induced transient growth stunting with partial recovery, and 500 µM significantly reduced biomass without affecting an unspecified additional trait. The dose- and compound-specific phenotyping establishes baseline toxicity thresholds usable for follow-up mechanistic work (e.g., uptake/transport assays) or risk assessment in legume crops.
bioRxiv · microbiologyConceptual
Bacteria have a fast pre-emptive shield against antibiotics, before they even bother changing their genes.
Antibiotic resistance in bacteria like Pseudomonas aeruginosa often depends on how easily drugs can get through the bacterial outer wall, through pore proteins called porins. Scientists discovered a small protein, PtrA, that makes bacteria resistant to a key antibiotic (imipenem, a carbapenem) not by getting rid of the pore, but by physically interfering with it to block drug entry. This response is triggered by metal signals (zinc and copper) and kicks in fast, acting as an early defense before the bacteria's slower strategy of shutting down pore production even begins. This reveals a previously hidden, quick-response layer of antibiotic resistance that could be a new target for drugs designed to keep bacteria vulnerable.
Technical view
PtrA is a small periplasmic protein identified as a regulator of OprD-dependent carbapenem permeability in Pseudomonas aeruginosa, conferring imipenem resistance without reducing OprD porin abundance, instead associating with OprD-containing membrane complexes to physically restrict drug entry. This mechanism is mechanistically distinct from and precedes the transcriptional CzcRS-mediated repression of oprD, and is triggered by zinc/copper as physiological signals, defining a two-phase adaptive resistance model: rapid post-translational periplasmic gating followed by slower transcriptional porin downregulation. PtrA represents a novel, non-transcriptional resistance node that could be targeted to resensitize P. aeruginosa to carbapenems, independent of existing porin-expression-focused resistance mechanisms.
bioRxiv · microbiologyBuildable
Southern Ocean microbes form distinct neighborhoods, and where you look changes what you find.
The Southern Ocean around Antarctica is full of microscopic marine life that forms the base of the food web, but most studies of it have focused on easy-to-reach coastal spots. This research instead sampled open ocean and sub-Antarctic waters, sequencing DNA (specifically a gene called 16S rDNA used to identify bacteria) from four different locations, and matched it with ocean measurements like temperature and currents. They found that microbial communities differed a lot between regions and even between nearby localities, shaped by both local water conditions and the physical difficulty of organisms spreading between distant sites. A few specific microbial groups showed up consistently everywhere, suggesting they are especially well-adapted to the harsh, varied conditions across this remote ocean region.
Technical view
The study used 16S rDNA high-throughput sequencing paired with oceanographic data to characterize marine microbial community composition across four localities spanning two Subantarctic sites (Magellan Strait, Beagle Channel) and two Antarctic open-sea regions (Eastern Indian, Central South Pacific). Results show significant inter-regional and inter-locality divergence in taxonomic composition, alpha diversity, and enriched taxa, attributable to both local environmental filtering and dispersal limitation, while taxa including Clade Ia, Amylibacter, and NS5/NS2b marine groups showed broader distribution across sites. The dataset extends Southern Ocean microbial biogeography beyond coastal-focused sampling and provides a baseline for modeling dispersal-vs-environment drivers of community assembly in under-sampled circumpolar and subantarctic waters.
bioRxiv · molecular biologyBuildable
Scientists sequenced the cactus bug that makes red dye, to find the genes behind its waxy coat.
The cochineal bug is a small insect known for its dense white waxy covering and, in a related species, for producing red dye, but it's also a pest that damages prickly pear cacti. This study built a full genetic blueprint (genome) of the wild cochineal bug for the first time, alongside data on which genes are active and which proteins are made. The researchers specifically hunted for a family of genes called fatty acyl reductases (FARs), which are known to help insects build their protective wax coatings, and found 26 of them, some of which appear to have multiplied and specialized within this species. This genetic map gives scientists the tools needed to understand — and potentially disrupt — how this pest builds its wax armor, which could inform pest control strategies for cactus farming.
Technical view
The authors report a 359-Mb de novo genome assembly for Dactylopius opuntiae alongside transcriptomic and proteomic data, identifying 26 fatty acyl reductase (FAR) genes central to epicuticular wax biosynthesis. Phylogenetic reconstruction across Coccoidea (scale insects) reveals lineage-specific tandem expansions of FAR genes in D. opuntiae, implicating gene duplication as a driver of this species' distinctive dense wax phenotype. This multi-omics resource enables comparative genomic studies of wax biosynthesis across scale insects and provides candidate gene targets for pest-control strategies targeting the insect's protective wax layer.
bioRxiv · molecular biologyBuildable
A cheap, reliable recipe for growing circular DNA rings used to edit genomes.
Scientists often need long loops of single-stranded DNA (a floppy, one-sided version of the usual double helix) as tools for editing genomes, building tiny DNA structures, or making diagnostic tests. These circular loops are tougher than straight DNA strands because cells' natural DNA-chewing enzymes have a harder time degrading a ring than a strand with loose ends, and they can be made much longer than DNA can be chemically synthesized. Until now, making them required expensive specialty kits or ordering custom DNA from a company. This paper lays out a step-by-step lab recipe using a virus-derived tool called a phagemid (basically hijacking a bacterial virus's DNA-copying machinery) to mass-produce pure circular DNA cheaply, making a widely useful material accessible to any lab.
Technical view
The authors present a scalable protocol for producing circular single-stranded DNA (cssDNA) using an M13 phagemid system, addressing the field's dependence on costly commercial synthesis or specialized reagents for genome-editing and nanotechnology applications. cssDNA's circularity confers exonuclease resistance and enables generation of long, sequence-defined constructs beyond chemical synthesis limits. The workflow appears optimized for reproducibility and yield at bench scale, likely detailing phage propagation, ssDNA extraction, and purification steps. Researchers building genome-editing templates (e.g., for HDR or prime editing donors) or DNA nanostructures could adopt this protocol directly to replace outsourced synthesis.
bioRxiv · molecular biologyBuildable
Researchers built a genome-wide GPS to catch DNA-unwinding enzymes in the act.
DNA helicases are molecular motors that unzip the double helix so it can be copied, repaired, or read — but finding exactly where in the genome they're working has been hard because their action leaves no visible trace. The researchers solved this by fusing a helicase to an enzyme that chemically tags any exposed single-stranded DNA, leaving tiny 'footprints' (mutations) wherever the helicase had unzipped DNA, which they then read out by sequencing the whole genome. Using this trick on a yeast helicase called Hrq1 (related to a human gene linked to disease), they discovered it works heavily at genes read by a specific gene-copying machine, especially the small genes that make transfer RNA. This gives scientists a general new method to map where any helicase acts across an entire genome at near-single-letter precision.
Technical view
The authors fuse DNA helicases to activation-induced cytidine deaminase (AID), which deaminates cytosines only in transiently exposed ssDNA, converting helicase unwinding events into strand-specific mutational footprints readable by whole-genome sequencing at near-nucleotide resolution. Applying this to S. cerevisiae Hrq1 (a RecQ4-family helicase and functional homolog of human RECQL4, implicated in Rothmund-Thomson syndrome) produced the first genome-wide activity map, revealing strong enrichment at RNA Pol III-transcribed loci, particularly tRNA genes, with strand bias favoring the non-template/transcript strand. This AID-fusion approach is generalizable to other helicases and translocases, offering a template for in vivo mapping of ssDNA-generating enzymes beyond ChIP-based methods.
bioRxiv · molecular biologyBuildable
A chemical tag on a signaling protein can silently rewire which partners it grabs.
Cells constantly send messages by attaching phosphate tags to proteins, and one common 'reader' of these tags is a protein module called an SH2 domain, which itself can also get tagged. This study asks what happens when the reader gets tagged too — does it change what the reader can read? Using a screening method (a modified dot blot, essentially a way to test many protein interactions on a membrane at once) with mutations that mimic permanent tagging, the researchers found that tagging one specific spot on a protein called PTPN11 changes its pickiness, making it bind fewer of its usual partners. This matters because PTPN11 and similar proteins are central hubs in cell growth signaling, and disrupting them is linked to cancer and developmental disorders, so understanding this extra layer of control could reveal new ways to intervene.
Technical view
SH2 domains recognize phosphotyrosine motifs to mediate signaling protein-protein interactions, but the regulatory effect of phosphorylation occurring within the SH2 domain itself (as opposed to on its ligand) was poorly characterized. The authors used phosphomimetic mutagenesis combined with a modified dot blot binding assay to interrogate two conserved intra-domain tyrosine phosphorylation sites across PTPN11-N, LYN, and SYK-C SH2 domains. They found that phosphomimicking Y63 in the PTPN11 N-terminal SH2 domain alters binding specificity, reducing affinity for physiologically relevant substrates — implicating this residue in an autoregulatory layer distinct from canonical ligand-based control. This establishes a scalable assay for probing SH2 phosphoregulation that could be extended to other domains implicated in RASopathies and cancer signaling.
bioRxiv · cell biologyBuildable
A new lab-grown cell line gives malaria researchers a working model of a key mosquito.
Anopheles stephensi is a mosquito species spreading malaria in cities, but scientists have had almost no lab tools — like cultured cells — to study its biology or engineer genetic controls against it, unlike better-studied mosquito species. This paper reports growing a new, self-sustaining line of cells taken from this mosquito's embryos, then verifying it's really from this species and figuring out its chromosomes. The team also tested different methods for getting foreign genetic material into these cells and found one reagent worked notably better than the common alternative, then used a light-producing reporter system to check which genetic 'on' switches work best in these cells. This toolkit fills a major gap, letting researchers now run genetic experiments directly in cells from this important disease-carrying mosquito instead of guessing from related species.
Technical view
The authors established and validated SDA-500, a novel embryo-derived cell line from Anopheles stephensi, an urban malaria vector for which molecular tools have been lacking. Species and karyotype were confirmed via mitochondrial COI barcoding and cytogenetics (revealing a diploid complement including a Y chromosome). They benchmarked transfection reagents, finding TransIT-PRO outperforms Lipofectamine-based methods, and used dual-luciferase reporter assays to evaluate promoter activity across candidates. This cell line provides a practical in vitro platform for functional genomics and testing genetic control constructs (e.g., gene drives) in A. stephensi, directly usable by vector biology labs for transfection-based screens.
bioRxiv · cell biologyConceptual
An ancient parasite's chromosome-grabbing protein reads a histone's bare tip, not its usual mark.
When cells divide, a molecular machine called the kinetochore grabs each chromosome and pulls it to the right place — normally this machine locks onto a special marker protein (CENP-A) built into the chromosome's DNA-packaging spool. But kinetoplastids, a group of ancient single-celled parasites, lack that marker entirely, so how their kinetochore finds the right spot has been a mystery. This study shows that one of their kinetochore proteins, KKT2, instead directly grabs the very tip of an ordinary packaging protein (histone H3) — and strikingly, even the smallest possible chemical modification to that tip completely blocks the grip. Using precise molecular measurement techniques, the researchers pin down exactly how this alternate recognition works, revealing a completely different solution evolution found for the same essential problem of finding chromosome attachment points.
Technical view
Kinetoplastids lack CENP-A, the centromere-specifying histone H3 variant universal to canonical eukaryotic kinetochores, raising the question of how their divergent kinetochore proteins achieve centromere-specific localization. The authors show via NMR spectroscopy and isothermal titration calorimetry that the centromere localization (CL) domain of Trypanosoma brucei KKT2 structurally resembles a ZZ domain and binds the free N-terminus of histone H3 directly, using an invariant aspartate conserved in known H3-binding ZZ domains. Critically, even N-terminal mono-methylation of H3 abolishes binding, indicating strict recognition of the unmodified free amino group — a mechanism orthogonal to canonical CENP-A-based recruitment. This suggests kinetoplastids evolved an alternative, modification-sensitive histone-reading strategy for centromere specification, offering a comparative model for kinetochore assembly diversity.
bioRxiv · cell biologyBuildable
Flipping a light switch inside cells raised a signaling molecule to fix heartbeats and Parkinson's tremors in mice.
Cells use a molecule called cAMP as an internal messenger to control things like heart rate and certain channels involved in Parkinson's disease. Instead of using drugs, which are blunt and slow, these researchers used optogenetics — inserting a light-sensitive enzyme into cells so that shining light on them precisely raises cAMP levels on demand. They showed that this light trigger sped up the beating of heart cells and, when aimed at a brain region involved in movement in mice with a Parkinson's-like condition, partly restored normal motor function and reduced abnormal spinning behavior. This proves that precisely controlling a single signaling molecule with light, rather than a drug hitting many targets, can meaningfully influence both heart rhythm and Parkinson's-like symptoms, pointing toward more targeted future therapies.
Technical view
The authors express a photoactivated adenylyl cyclase (PAC(S27A)) to optogenetically raise intracellular cAMP with light, then examine downstream effects on HCN channels, which are cAMP-gated and implicated in both cardiac pacemaking and PD pathophysiology. Light-induced cAMP elevation activated HCN4 to increase cardiomyocyte beating rate, while unilateral PAC(S27A) expression in the substantia nigra pars compacta of mice produced light-dependent rotational behavior attenuable by HCN inhibitors, and partially rescued motor deficits in an MPTP-induced PD model with concurrent HCN2 changes. This establishes PAC(S27A) as a tractable optogenetic tool for dissecting cAMP-HCN signaling causally in both cardiac and dopaminergic circuits, with potential application toward HCN-targeted PD or arrhythmia interventions.
bioRxiv · developmental biologyConceptual
IVF and embryo freezing leave a mitochondrial scar on the heart that lasts into adulthood.
IVF (in vitro fertilization) has helped create over 10 million babies, and children conceived this way sometimes show subtle heart differences like altered heart structure and higher blood pressure, but nobody knew why. This study looked at mitochondria — the energy-generating parts of cells — in mouse embryos made by IVF, including ones that were frozen and thawed (vitrified), and tracked what happened as those mice grew into adults. They found that IVF disrupted the embryos' mitochondrial energy processing and their balance of reactive, potentially damaging molecules right from the earliest stage, and traced these disruptions forward to find they persisted in the adult heart. This is the first evidence linking early embryo-stage mitochondrial stress from fertility treatments directly to lasting heart effects in adulthood, suggesting doctors may need to consider long-term cardiovascular monitoring for IVF-conceived individuals.
Technical view
Using the IGS-CD1 mouse model, the authors compared blastocysts derived from natural mating versus IVF (transferred fresh or after vitrification-warming), assessing mitochondrial redox balance and metabolic function at the blastocyst stage and then following offspring hearts into adulthood. IVF reduced total blastocyst mitochondrial content/function and altered redox status, and these early perturbations were traceable to persistent cardiac abnormalities in adult offspring, providing a mechanistic link between preimplantation ART exposure and previously reported ART-associated cardiovascular phenotypes (cardiac remodeling, elevated blood pressure). This is apparently the first study to directly connect blastocyst-stage mitochondrial dysfunction to adult cardiac outcomes in an ART model, offering a framework for testing interventions (e.g., antioxidant supplementation during IVF) to mitigate long-term cardiovascular risk.
bioRxiv · developmental biologyConceptual
Migrating cells physically sculpt the very tissue they need in order to migrate — a two-way construction crew.
As a vertebrate embryo's head forms, two things happen almost simultaneously: neural crest cells crawl away from the developing brain tube, and that tube folds itself closed (neurulation). Scientists used to think these were separate, independent programs, but this study shows they're actually in constant physical conversation. As the migrating neural crest cells crawl out, they remodel a scaffold protein called fibronectin in the space between tissues, and this remodeled scaffold both keeps the tissues properly separated and helps the neural tube fold correctly by letting its cells rearrange and squeeze into shape. This remodeling depends on a molecular pair of scissors called MMP14 that only the migrating cells carry, revealing that cell migration and tissue folding aren't separate assembly lines but a feedback loop where each shapes the other — insight relevant to birth defects like neural tube closure failures.
Technical view
The authors demonstrate that cephalic neural crest migration and neural tube closure, long treated as tissue-autonomous processes, are mechanochemically coupled via extracellular matrix remodeling. As neural crest cells delaminate and invade adjacent mesoderm, they remodel fibronectin at the neural crest-neural plate boundary in an MMP14 (membrane-bound metalloproteinase)-dependent manner, generating an ECM interface that both physically separates the tissues and permits radial intercalation and apical constriction in the neural plate necessary for tube closure. Loss of neural crest-specific MMP14 or migration disrupts this ECM remodeling and impairs neurulation, establishing collective cell migration as a mechanical prerequisite for adjacent morphogenesis. This reciprocal feedback model reframes neurulation defects (a major class of human birth defects) as potentially arising from disrupted neural crest-ECM crosstalk rather than neural plate-intrinsic failure alone.
bioRxiv · developmental biologyConceptual
A surprise gene helps build the gut's tiny pipes that pump fat into your blood.
Deep inside the lining of your intestine are microscopic lymph vessels called lacteals that soak up fat from digested food, and they only work because tiny muscle cells around them squeeze rhythmically to push the fluid along. Scientists didn't know how the different support cells building this muscle-vessel unit actually talk to each other during development. Using single-cell gene mapping, mouse genetic tracing, and lipid-absorption tests, this team found that a gene called Notch3 acts like a foreman, coordinating separate crews of connective-tissue cells so the muscle wrapping forms correctly. It matters because when this assembly goes wrong, it can cause lymphatic diseases that are currently very hard to treat.
Technical view
The study uses developmental single-cell RNA profiling, Cre-based lineage tracing, and conditional mouse genetics to dissect assembly of the muscular-lacteal complex (MLC) in intestinal villi. Notch3 is identified as a key coordinator between distinct mesenchymal lineages, promoting smooth muscle differentiation within the PDGFRα lineage, even though PDGFRβ lineage cells themselves do not directly contribute to villus smooth muscle — implying a non-cell-autonomous signaling relay. Functional lipid-absorption assays link MLC integrity to physiological lacteal pumping. This establishes a genetic entry point for dissecting mesenchymal crosstalk in lymphatic vessel maturation, relevant to lymphatic dysfunction disorders.
bioRxiv · ecologyConceptual
Sulfur trapped in ancient hunted bones maps how Stone Age people roamed Spain.
Long before writing, hunter-gatherers in northern Spain moved across the landscape following game and resources, but figuring out exactly how far and how often is hard when all you have is old bones. Researchers measured sulfur, carbon, and nitrogen isotopes — chemical signatures locked into bone collagen that vary depending on the specific soil and region an animal grazed in — from over 900 hunted animal bones spanning nearly 100,000 years. By combining these isotope 'fingerprints' with protein analysis, climate reconstruction, and statistical dating models, they could trace which regions the meat (and therefore the hunters) came from over time. This gives a rare, detailed picture of how prehistoric human mobility patterns shifted with climate and culture, from Neanderthals through early modern humans.
Technical view
The team applied δ34S isotope analysis, cross-validated with δ13C and δ15N and palaeoproteomic species identification, to 905 anthropogenically modified ungulate bone collagen samples from 16 Cantabrian sites spanning MIS 5 to 1 (100–7 ka BP). Sulfur isotopes serve as a geolocation proxy since bedrock-derived sulfate signatures vary spatially, enabling isoscape mapping of animal (and by extension human) provenance. Combined with Bayesian chronological modeling and palaeoclimatic reconstruction, this reconstructs shifts in hunter-gatherer territorial ranges and resource exploitation strategies across the Mousterian-to-Mesolithic transition. The approach offers a replicable isotopic-provenancing framework for other regions with dense faunal assemblages.
bioRxiv · ecologyBuildable
AI that learns population biology AND makes the best call for saving an endangered fish.
When trying to save an endangered species, wildlife managers face a tough tradeoff: models detailed enough to reflect real biology are often too complex to use for picking the best action, while simple decision models ignore important biological nuance. This research fuses two AI approaches — one that builds a rich, data-grounded picture of how a population actually grows and shrinks (an integrated population model), and another that's very good at learning optimal strategies under uncertainty (deep reinforcement learning, the technique behind game-playing AI). Together they let managers get recommendations that are both ecologically realistic and genuinely optimized. They tested this on real conservation efforts to help the endangered Rio Grande silvery minnow, showing it can guide decisions like when and how many fish to release into the wild.
Technical view
The framework couples integrated population models (IPMs), which synthesize multiple data streams (survival, reproduction, abundance surveys) into a unified demographic model, with deep reinforcement learning (DRL) to solve for adaptive management policies over high-dimensional, uncertain state spaces — avoiding the simplification typically required by classical stochastic dynamic programming. Applied to the Rio Grande silvery minnow supplementation program, the IPM-DRL pipeline learns policies directly from the fitted ecological model rather than a reduced-form approximation. This is a template for closing the gap between ecological realism and decision-theoretic optimality in endangered species management, and could generalize to other supplementation or harvest-control problems where population models already exist.
bioRxiv · ecologyRunnable
A no-code app helps scientists double-check what AI 'heard' in wildlife audio recordings.
Researchers now use AI to sift through huge amounts of field audio recordings to detect specific animal calls, but the AI often makes mistakes, so humans still need to listen and confirm each detection before trusting the results. Until now that verification step was done through messy, ad hoc spreadsheets and manual processes prone to error and hard to document. PAMalytics is a free, open-source app that runs right in your web browser (no programming needed) to organize and standardize this human review process. It matters because reliable species monitoring — for conservation, environmental impact assessments, and biodiversity tracking — depends on trustworthy validation, not just a good classifier.
Technical view
PAMalytics is an open-source, local browser-based, no-code application targeting the post-classification validation bottleneck in passive acoustic monitoring (PAM) pipelines. It provides structured workflows for reviewers to confirm or reject automated species-classifier detections, addressing the gap between classifier output and downstream occupancy/detection models that assume validated data. By standardizing and logging the validation process, it improves transparency, reduces transcription/consolidation errors, and creates auditable records — useful for anyone building PAM-based biodiversity monitoring or regulatory reporting pipelines who currently relies on manual spreadsheet review.
bioRxiv · biochemistryBuildable
Same enzyme family, three different 3D shapes, three different jobs cutting TB's cell wall sugar.
Mycobacteria — the family that includes tuberculosis — build their tough cell walls using an unusual sugar chain called arabinan, and figuring out how to break it down could help fight these bacteria or study their biology. A gut bacterium was found that can fully dismantle this sugar using a toolkit of enzymes, including three related ones (called GH172) that snip the chain from its ends. This study shows that even though these three enzymes look similar, they cut the chain at different specific points and, surprisingly, assemble themselves into completely different 3D shapes. Using precisely designed test sugars, custom chemical probes, and powerful imaging (X-ray and cryo-electron microscopy), the researchers mapped out exactly how form relates to function — insight useful for designing drugs or biotech tools that target this cell-wall sugar.
Technical view
Three GH172 exo-α-arabinofuranosidases from Dysgonomonas gadei, previously shown to fully degrade mycobacterial arabinan (a component of arabinogalactan and lipoarabinomannan), were characterized against defined synthetic substrates and shown to have distinct linkage specificities. The team developed α-arabinofuranosyl cyclophellitol aziridine-based covalent inhibitors and activity-based probes to interrogate active-site chemistry, and solved X-ray crystal and cryo-EM structures revealing markedly different quaternary assemblies among the three homologues despite shared catalytic fold. This links oligomeric architecture to functional divergence within a single GH family and provides new chemical tool compounds (activity-based probes/inhibitors) usable for studying arabinan-processing enzymes or mycobacterial cell-wall biology more broadly.
bioRxiv · bioengineeringBuildable
One round of lab sorting turned mouse antibodies into human-ready Alzheimer's drug candidates.
To make an antibody drug from an animal immune response usable in humans, scientists must 'humanize' it — reshape it to look human enough that the immune system won't reject it — while still binding its target tightly, which traditionally takes many slow rounds of trial and error. Here, researchers targeting the amyloid-beta protein linked to Alzheimer's disease combined an efficient antibody-display technique with a smarter, structure-guided way of choosing which parts of the antibody to humanize. They immunized cells, screened billions of antibody variants using fluorescence sorting, and speeply humanized two promising candidates in a single screening round instead of many. This dramatically speeds up an otherwise slow, expensive bottleneck in developing new antibody therapies, here demonstrated for a major Alzheimer's drug target.
Technical view
The workflow combines immune yeast Fab-display (library diversity ~3.5×10^8) with magnetic enrichment and FACS to isolate anti-Aβ1-42 aggregate-reactive clones, then applies structure-guided single-round focused humanization by sampling framework positions predicted to support CDR conformation or VH/VL domain packing, rather than iterative loop-by-loop humanization. Two lead clones, CLAB17 and CLAB45, were converted to full IgG and advanced through this pipeline, achieving binding-positive humanized candidates from a single FACS sorting round. This demonstrates a generalizable, faster alternative to conventional multi-round CDR-grafting/back-mutation humanization for early-stage antibody developability assessment, applicable beyond amyloid-β targets.
bioRxiv · bioengineeringRunnable
A light-based blood-flow sensor gets confused by 'boring' tissue — here's how much it skews results.
Doctors increasingly use a laser-light technique called diffuse correlation spectroscopy to non-invasively track blood flow in tissue, by shining light in and watching how the pattern flickers as red blood cells move. But tissue isn't just moving blood cells — it also contains completely still material and slowly-shifting material, both of which distort the flicker pattern and can throw off the blood-flow reading. This study built physical models ('phantoms' — like gel blocks with a tube of flowing liquid standing in for blood vessels) to measure exactly how much these non-blood components skew the results. Understanding and correcting for this matters because it affects how accurately doctors can trust these devices for monitoring things like brain or muscle blood flow in real patients.
Technical view
Using continuous-wave diffuse correlation spectroscopy (cw-DCS), the authors quantify how static and slow-dynamic scatterer populations bias the derived blood flow index (BFI), which is conventionally computed assuming decorrelation is dominated solely by fast-moving red blood cells. Agar-based phantoms embedded with a flow tube were used to isolate and measure the fractional contribution of each scatterer class (static, slow-dynamic, fast-dynamic) to the measured autocorrelation decay rate. Results show static/slow components meaningfully alter derived BFI values, implying that current single-component fitting models used in clinical cw-DCS devices may need multi-component correction terms for accurate absolute (not just relative) flow quantification — directly relevant to anyone calibrating or validating DCS hardware.
bioRxiv · bioinformaticsRunnable
Shrinking a protein-AI model to save memory can secretly break it on the one case that matters.
Big AI models that predict how mutations affect proteins are often 'compressed' (quantized) to run faster and cheaper, using lower-precision numbers instead of full precision, and this is usually judged safe by checking if average accuracy across many tests barely moves. This study tested six different compression settings on a popular protein-AI model (ESM-2) across 201 real mutation-effect benchmarks covering 2.4 million variants. The catch: even when the average accuracy looked essentially unchanged, one specific compression method silently wrecked performance on an individual test, cutting its accuracy score by more than half. This means relying on average benchmark scores to decide if a compressed AI model is 'safe' to deploy can hide serious, deployment-breaking failures on important individual cases.
Technical view
The authors benchmark six numerical precision configurations of ESM-2 (650M–15B parameters) on bulk embedding extraction and deep mutational scanning (DMS) variant-effect scoring, evaluated across the full ProteinGym substitution benchmark (201 assays, 2.41M variants) with paired bootstrap clustering by protein. While no configuration shifts mean Spearman correlation by more than 0.007 at any scale, INT8 dynamic quantization — statistically indistinguishable from fp32 on the mean at 3B parameters (p=0.34) — collapses one specific assay's correlation from ρ=0.591 to 0.223. This demonstrates that mean-based benchmark evaluation is an unreliable criterion for quantization deployment decisions in protein language models, and practitioners should instead audit worst-case, per-assay degradation before shipping quantized variant-effect predictors.
bioRxiv · biophysicsBuildable
Scientists made proteins glue themselves together permanently, no enzymes needed.
Some bacteria build super-tough surface fibers by having two amino acids in a protein spontaneously fuse into a permanent chemical bond, no external glue or enzyme required. Here researchers designed brand-new proteins from scratch that do this same self-stitching trick, both within a single protein chain and between two separate ones. They even split the design into two halves that only snap together and lock when mixed, and tuned it so temperature controls exactly when the bond forms. The payoff is being able to build large, rigid, precisely shaped molecular structures — like a 215,000-atomic-mass-unit ring — that are essentially welded together and won't fall apart, useful for building durable nanoscale materials or vaccine scaffolds.
Technical view
The authors computationally designed de novo proteins that autocatalytically form intramolecular and intermolecular isopeptide bonds (side-chain amide linkages, as seen in Gram-positive pilin proteins), producing 50+ validated designs confirmed by mass spec and 5 crystal structures. They engineered split constructs that crosslink upon combination, orthogonal to the widely used SpyTag/SpyCatcher system, with bond formation kinetics tunable by temperature. These modules were used to rigidly crosslink domains into large (up to 215 kDa) symmetric, covalently locked ring assemblies — a generalizable toolkit for programmable, irreversible protein nanoarchitecture beyond existing peptide-tag crosslinking systems.
bioRxiv · cancer biologyConceptual
Cancer cells hijack a metabolic enzyme by tagging and pairing it up for extra fuel.
Cells need a molecule called NAD+ to run their metabolism, and an enzyme called NAMPT is the key factory that makes it. This study found that several cancer-driving proteins (kinases that are often mutated or overactive in tumors) directly chemically tag NAMPT at one specific spot, and this tag works together with NAMPT pairing up with a copy of itself to switch the enzyme into overdrive. When researchers blocked either the tagging or the pairing, cancer cells made less NAD+, grew slower, and formed fewer colonies. This matters because it reveals a druggable weak point — cutting off this tag-and-pair switch could starve cancer cells of the fuel-making machinery they depend on.
Technical view
NAMPT, the rate-limiting NAD+ salvage-pathway enzyme, is shown to be a direct phosphorylation substrate of oncogenic tyrosine kinases (ALK, INSR, IGF1R, PDGFRA), with phosphoproteomics identifying Y188 as the dominant site, notably downstream of the NPM1::ALK fusion. Y188 phosphorylation and NAMPT dimerization act cooperatively to boost catalytic activity and NMN/NAD biosynthesis; a Y188F mutant or dimerization-disrupting mutation each independently impair enzymatic activity, proliferation, and clonogenic growth. This defines a kinase-NAMPT signaling axis as a candidate therapeutic target for NAD+-dependent tumors, with Y188 phosphorylation or dimer-interface disruption as potential intervention points.
bioRxiv · cancer biologyBuildable
AI microscope hunts for one weird hybrid cancer cell among millions of normal blood cells.
When cancer spreads, some rare cells appear to be hybrids — part tumor, part immune cell — floating in the bloodstream, and finding these needle-in-a-haystack cells under a microscope is extremely hard because they're so sparse and every sample looks different. The researchers built a two-step system: first they use each animal's own unstained blood sample as a personalized baseline to filter out background noise, then they train a neural network (a pattern-recognizing AI) on labeled cell images to flag genuine candidates. They tested this in mice with pancreatic tumors, imaging blood cells tagged with fluorescent markers for tumor and immune cell identity. This kind of automated detection could make it far more practical to track cancer spread in blood samples instead of painful tissue biopsies.
Technical view
The authors present a two-stage pipeline for detecting rare ECAD+/CD45+ circulating hybrid cells (CHCs) in mouse PBMC preparations via multichannel fluorescence microscopy: first, animal-specific unstained control images establish per-mouse background distributions for candidate enrichment, then a CNN trained on blinded multi-annotator consensus labels (DAPI, ECAD, CD45 crops) classifies candidates. Validated across 28 mice (tumor-bearing and tumor-naive), this specimen-normalized enrichment-plus-classification approach addresses the core challenge of rare-cell detection amid inter-sample background variability, and the framework could generalize to other rare circulating cell types with adaptation of channel inputs and training labels.
bioRxiv · cancer biologyRunnable
A newly found drug tricks two proteins into teaming up to kill mutant mast cell cancer.
Certain mast cell cancers are driven almost entirely by one specific mutation in a gene called KIT, and this study looked for a drug that kills only cells carrying that mutation while sparing normal cells. Using cells grown from stem cells of actual patients, the team screened a large library of drugs and found one, called LDC3416, that selectively kills the mutant cells. Digging into how it works, they discovered it acts as a 'molecular glue' — essentially super-gluing two proteins (PDE3A and SLFN12) together into a cell-killing complex, and the mutant cancer cells happen to make extra amounts of both proteins, making them uniquely vulnerable to this glue effect. This offers a promising, more targeted treatment strategy for a disease that currently has few options.
Technical view
Using KIT D816V patient-iPSC-derived hematopoietic and mast cell models, the authors screened FDA-approved and experimental compounds and identified LDC3416 as selectively cytotoxic to KIT D816V-mutant cells across multiple lineages. Mechanistic profiling revealed LDC3416 acts through the PDE3A-SLFN12 molecular glue pathway, and that KIT D816V signaling upregulates both PDE3A and SLFN12 expression, conferring selective sensitivity to glue-induced complex formation and cell death. This establishes molecular-glue induction of the PDE3A-SLFN12 axis as a mutation-selective therapeutic strategy for clonal mast cell malignancies, with LDC3416 as a lead compound for further preclinical development.
bioRxiv · systems biologyConceptual
Losing genes can actually help bacteria — if the gene was expensive to keep making.
Every gene a bacterium carries costs it energy and resources to constantly manufacture into protein, so scientists wondered whether ditching some genes could actually free up resources and boost growth. Using E. coli as the test subject, they sorted genes into categories — essential, important, mostly-neutral, or even fitness-boosting when deleted — and combined this with data on how much of the cell's protein-making budget each gene consumes. They found that expensive-to-produce genes are more likely to matter for fitness, but surprisingly, deleting genes rarely helps growth simply by freeing up that production budget. This nuances the popular idea that bacteria are constantly trimming costly, useless genes to save energy — the real picture of why gene loss helps is more complicated.
Technical view
The study integrates fitness measurements with proteomic abundance data across the E. coli genome to test whether gene loss improves fitness primarily by reducing proteome allocation cost. Genes are classified into essential, important, mean-effect, and fitness-enhancing categories; mean-effect genes constitute 31-75% of the total proteome mass fraction depending on growth condition (highest in LB medium), with core mean-effect genes enriched for transmembrane transport functions. Comparison against genome-scale metabolic models with proteome constraints (ME-models) shows that while high proteomic-cost genes are more likely to influence fitness, fitness-enhancing deletions rarely act via simple resource-reallocation savings — implying deletion benefits arise from more complex regulatory or metabolic effects than proteome-burden relief alone.
bioRxiv · neuroscienceConceptual
Staying awake longer makes your brain fluid pulse harder to flush out waste.
While you sleep, slow rhythmic pulses in your brain's blood vessels help pump cerebrospinal fluid (the liquid cushioning your brain) through brain tissue, flushing out waste — a process called brain clearance. This study kept healthy people awake for 34 hours straight while continuously scanning their brains and measuring alertness, to see what drives these fluid pulses beyond just sleep itself. They found the pulses got stronger the longer people stayed awake and also naturally peaked in the early morning regardless of sleep loss, and this tracked closely with alertness levels and activity in brain circuits that keep us awake — but not with a stress hormone or with typical sleep-related brain wave patterns. This suggests the brain's built-in waking and body-clock systems, not just sleep itself, help regulate how well it cleans itself, which matters for understanding conditions linked to poor sleep and waste buildup, like some neurodegenerative diseases.
Technical view
In a 34-hour sleep deprivation protocol with dense longitudinal fMRI/EEG and NIRS sampling in healthy adults, the authors quantified cerebrospinal fluid low-frequency oscillations (LFOs, the vasomotion-driven signal thought to propel CSF-based brain clearance) and dissociated circadian versus homeostatic sleep-pressure contributions. CSF LFO amplitude increased monotonically with time awake and showed independent diurnal modulation peaking in early morning; LFOs correlated strongly with vigilance measures and reticular activating system (wake-promoting brainstem/thalamic) activity, but not with plasma norepinephrine or EEG slow-wave activity. The findings implicate arousal-circuit and circadian signaling, rather than adrenergic tone or classic sleep EEG markers, as regulators of vasomotor-driven CSF dynamics — relevant to models linking sleep loss, arousal state, and glymphatic-type clearance dysfunction.
bioRxiv · molecular biologyConceptual
A protein tied to premature-aging disease jams the cell's DNA repair crew.
Hutchinson-Gilford Progeria Syndrome is a devastating rare disease that makes children age dramatically fast and die young, caused by a mutation that produces a toxic, truncated protein called progerin (a warped version of a normal structural protein in the cell's nucleus). Cells with progerin build up broken DNA strands and seem worse at fixing them, so this study investigates exactly how progerin interferes with the two main repair systems cells use to mend broken DNA — one precise (using a matching template) and one quicker but sloppier. Using cells engineered to make progerin, the researchers show that break repair is impaired, and start dissecting which repair pathway suffers and why. Understanding this could explain how progerin drives accelerated aging and possibly shed light on ordinary aging too, since everyone makes small amounts of progerin naturally.
Technical view
HGPS results from a LMNA point mutation activating a cryptic splice site, producing farnesylated truncated lamin A ('progerin'), which is associated with elevated genomic double-strand break (DSB) burden and altered DSB repair balance between homologous recombination (HR, template-directed and accurate) and end-joining (EJ, error-prone). The authors use progerin-expressing cell models to directly assess DSB repair kinetics and pathway usage, finding that break healing is impeded in progerin-expressing cells. This work aims to mechanistically link progerin expression to specific repair pathway defects, with implications for both HGPS pathology and physiological aging, since low-level progerin production occurs in normal cells as well.
bioRxiv · molecular biologyBuildable
A new method builds dozens of custom DNA sequences at once from one cheap mixed batch.
Making custom DNA in the lab is expensive when done piece by piece, but far cheaper when you order many short DNA snippets together as a mixed 'pool' — the catch is sorting that jumbled pool back into the exact separate DNA products you actually want. This paper introduces a method called Oligo Pool Sidewinder assembly that solves this sorting problem: it uses a computer program to design each short snippet with built-in matching rules, so that in a single test tube reaction, hundreds of pieces spontaneously find their correct partners and stitch together correctly into dozens of distinct final DNA constructs. This makes large-scale, low-cost DNA construction from cheap pooled snippets practical, which matters for anyone building genes, gene circuits, or other custom DNA at scale — from researchers to biotech companies.
Technical view
The paper introduces Oligo Pool Sidewinder assembly, a method combining a computational bespoke-oligo design workflow with a defined set of construction rules to enable one-pot, parallel, highly multiplexed assembly of hundreds of DNA fragments from unpartitioned synthetic oligo pools into dozens of distinct, correctly-formed double-stranded DNA constructs simultaneously. This directly addresses the yield/fidelity tradeoff inherent to pooled oligo synthesis (cheap but unindexed) versus individually synthesized oligos (isolatable but costly), offering a scalable route to reduce cost, labor, and turnaround for de novo multi-construct DNA production, with the computational design component likely adaptable to other multiplexed assembly schemes.
bioRxiv · cell biologyBuildable
A new lab trick reveals how blood stem cells retune their protein-building machinery as they mature.
Every cell uses molecules called tRNAs to translate genetic instructions into proteins, like adapters that match each three-letter DNA code to the right amino-acid building block. Scientists didn't know if these adapters change as a blood stem cell decides to become, say, a red blood cell or an immune cell. Here researchers built a new tool that reads both a cell's tRNAs and its regular gene activity at the same time, in thousands of individual cells from human bone marrow. They found specific tRNA patterns line up with 'stemness' versus different mature blood cell fates, revealing a whole extra layer of cellular identity nobody had mapped before. This matters because it suggests cells fine-tune protein production, not just gene activity, to become who they are.
Technical view
The authors developed sc-STM-seq, a scalable single-cell method that co-profiles tRNA abundance/splicing and mRNA transcriptomes in the same cell, addressing the technical difficulty of capturing short, heavily modified tRNAs alongside standard scRNA-seq. Applied to human bone marrow, they construct differentiation trajectories and correlate specific tRNA isodecoder expression and splicing events with stemness and lineage commitment. The result is a hematopoietic tRNA atlas positioning tRNA regulation as an independent axis of cellular heterogeneity beyond transcript-level control. Practitioners in translational control or codon-usage optimization could apply sc-STM-seq to other differentiation systems or mine this dataset for lineage-specific tRNA biomarkers.
bioRxiv · cell biologyBuildable
A new tool cracks open the black box of AI that paints fake microscope images of cell organelles.
In silico labeling is a trick where an AI looks at a plain, unstained microscope image and predicts where specific cell structures like mitochondria would glow if you'd stained them, saving time and chemicals. The problem is nobody could tell if the AI was truly 'seeing' real biology or just guessing from unrelated visual patterns, which makes it risky to trust. The researchers built Mask Interpreter, a method that shows exactly which visual features the AI relies on for each organelle, like a fingerprint of its reasoning. It turns out the models genuinely key off real biological structures, and the tool can also catch experimental glitches ('batch effects') and flag likely mistakes even without a ground-truth image to compare against. This matters because it makes a powerful but opaque prediction technology safe enough to actually trust in biology labs.
Technical view
Mask Interpreter is a semantic visual interpretability framework for image-to-image translation models used in label-free organelle prediction, generating organelle-specific 'explanation signatures' rather than generic saliency maps. Benchmarked against standard explainable-AI (xAI) methods, it better discriminates authentic biological signal from spurious correlations the network may exploit, and additionally functions as a diagnostic for batch effects and localized prediction error even absent paired fluorescence ground truth. This positions it as both a validation layer for deploying in silico labeling pipelines and a QC tool for microscopy datasets. Groups building or auditing cross-modality translation models could adopt the explanation-signature approach as a drop-in interpretability check.
bioRxiv · cell biologyConceptual
Cells have a fast-acting 'reset button' for their internal skeleton, and scientists just found its safety switch.
Cells contain a scaffold of protein filaments called actin that reshapes itself constantly to change a cell's shape or heal a wound. When calcium floods into a cell, a protein called INF2 triggers a rapid, sweeping rebuild of this scaffold, but too much INF2 activity is linked to kidney and nerve diseases, so it needs strict control. Using high-resolution microscopy that tracks single molecules plus biochemistry and structural analysis, the researchers found two built-in brakes: one where INF2 folds up on itself, and another where its tip grabs the side of actin filaments to stop them growing too long. Breaking this second brake keeps INF2 switched on too long and disrupts the cell's outer membrane. This matters because it explains, at a molecular level, how cells prevent a normal repair process from spiraling into disease.
Technical view
The study dissects the regulatory circuit governing INF2, the formin responsible for the Calcium-mediated Actin Reset (CaAR) reaction, combining live-cell imaging, single-molecule tracking, biochemistry, and structural analysis. Beyond the known intramolecular autoinhibition, they identify a second mechanism in which the INF2 N-terminus binds laterally to actin filament sides, capping elongation and promoting reversion to the autoinhibited state, a negative feedback loop. Disrupting this side-binding interaction prolongs INF2 activity and perturbs plasma membrane organization, linking the mechanism directly to the kidney and neuronal pathologies associated with INF2 hyperactivity. This gives structural and cell biologists a specific interface (N-terminus-filament side contact) as a candidate target for modulating formin activity therapeutically.
bioRxiv · biochemistryBuildable
A drug tool meant to study one enzyme turned out to be hitting an Alzheimer's-linked protein too.
PLD3 is a protein tied to Alzheimer's disease and immune signaling, but it only works once other enzymes chop it into its active form, and nobody knew which enzymes did the chopping. Researchers used a chemical called E64d, known to block a family of protein-cutting enzymes called cysteine cathepsins, to investigate. To make sure E64d was doing what everyone assumed, they built a modified 'tagged' version of it and used a technique that maps every protein a drug actually sticks to in a cell. Surprisingly, E64d also grabbed onto several other proteins beyond cathepsins, including ones linked to detox pathways and gene expression, while also pointing to which cathepsins process PLD3. This matters because it both clarifies a mechanism relevant to Alzheimer's and warns researchers to double-check what their 'specific' inhibitors are really doing.
Technical view
The authors use activity-based protein profiling (ABPP) with a propargyl-alkyne analogue of the cysteine cathepsin inhibitor E64d to map its full covalent target landscape in cells, going beyond its canonical cathepsin targets. The chemoproteomic survey uncovers unexpected covalent engagement with bleomycin hydrolase (BLMH), KEAP1, SPT5 (SUPT5H), and an asparagine synthetase, indicating substantial off-target reactivity for a widely used cathepsin probe. In parallel, they use this pharmacological toolkit to implicate specific cysteine cathepsins in the proteolytic maturation of PLD3, a 5'-3' exonuclease relevant to Alzheimer's disease and innate immune signaling. This is directly useful to chemical biologists as a target-engagement dataset for reinterpreting past E64d-based studies and for cathepsin-selective probe design.
bioRxiv · biochemistryConceptual
Cryo-electron microscopy caught a bacterial potassium pump in the act of being switched off.
Cells need to carefully balance potassium and hydrogen ions to survive, a job done by proteins called cation-proton antiporters that swap one ion for the other across the cell membrane. Researchers studied one such pump in E. coli, called YcgO, which trades potassium for protons, and a regulatory protein called PtsN that can shut it down. Using cryo-electron microscopy, a technique that flash-freezes molecules and images them with electrons to reveal their 3D shape, they captured detailed snapshots of YcgO both alone and locked by PtsN. The structures show YcgO is a two-part (dimeric) machine with extra internal 'sensor' domains that PtsN grabs to jam the pump's moving parts. This matters because it reveals a previously unclear architecture and control mechanism for a whole class of potassium transporters found across many organisms.
Technical view
Using cryo-EM at 3.4 Å and 3.2 Å resolution, the authors solve structures of the E. coli CPA1 antiporter YcgO in its K+-bound occluded state and in complex with the unphosphorylated regulatory protein PtsN. YcgO is a homodimer with appended cytosolic RCK and CorC domains that couple ion-binding events to conformational changes in the transport helices; PtsN docks onto the CorC domains to allosterically inhibit helix movement and thus efflux. This provides a structural rationale for phosphorylation-state-dependent regulation of a K+-specific CPA1 transporter via a PTS-system component, a regulatory logic distinct from previously characterized Na+/H+ antiporters. Structural biologists working on ion transport or bacterial signaling could use these coordinates to model analogous RCK/CorC-domain regulatory interfaces in other CPA-family transporters.