2026-08-19 · IST

Wednesday, 19 August 2026

299 new items across 6 fields, each explained in plain words. Jump to a section:

AI

AI & Machine Learning

46 new
arXiv · cs.ROConceptual★ flagship

Don't Drop the BATON: Long-Horizon Robot Manipulation via Agentic Subtask Exploration and Transition-aware Memory

A language-model 'director' helps robots chain many delicate steps without errors piling up.

Robots are getting good at single manipulation skills — say, twisting a cap or inserting a plug — but real chores string dozens of these together, and a tiny mistake early on snowballs until the whole task collapses. This work uses a smart division of labor: a language-model 'agent' acts as director, planning the steps in words and moving the arm through empty space with simple geometry, while calling in a specialized hand-eye skill model only for the tricky touch-heavy moments. The clever bit (called BATON) is teaching the system to explore one subtask at a time instead of the whole task at once — otherwise the practice needed explodes exponentially with each added step — and to remember not just how a step ended but what the next step needs to begin cleanly. It also keeps a written 'memory' of what worked so it can adapt. This matters because reliable multi-step manipulation is the gap between lab demos and robots that actually do useful sequences of chores.

Technical view

BATON keeps a frozen VLA policy for contact-rich segments and puts an LLM agent in charge of language-level planning, analytic free-space motion primitives, and language-memory adaptation. Its two contributions target the failure modes of naive whole-task test-time exploration: it decomposes exploration to the subtask level (avoiding the ~T^K episode cost and the credit-assignment ambiguity of stage-agnostic failures) and it adds transition-aware memory that represents both exit and entry conditions between primitives, so one subtask's outcome is checked against the next's preconditions rather than silently constraining it. Practitioners could adopt the frozen-VLA-plus-LLM-orchestrator pattern and the entry/exit transition contract to compose existing skill policies into longer horizons without retraining. The gain is sample efficiency and diagnosable failures in K-stage manipulation.

arXiv · cs.LGRunnable★ flagship

Q-based Variational Inverse Reinforcement Learning

Teach AI human preferences by watching experts — and know how confident it should be.

We often can't write down exactly what we want an AI to value, but we can show it examples of good behavior. Inverse reinforcement learning flips normal learning around: instead of being told the reward, the system watches an expert and infers what reward would explain those choices. This method, QVIRL, does that in a Bayesian way — meaning it doesn't just guess one reward function but produces a whole spread of plausible ones, so it can say 'I'm unsure here.' It gets there by learning a probability distribution over the quality-scores of actions (the 'Q-values') rather than over rewards directly, which keeps it scalable. That uncertainty is valuable for safety and for actively asking humans about the cases it's least sure of.

Technical view

QVIRL is a Bayesian IRL method that recovers a posterior over reward functions by learning a variational distribution over optimal Q-values rather than parameterizing rewards directly, which sidesteps expensive inner-loop MDP solving and improves scalability. The Q-space formulation yields calibrated uncertainty alongside point performance, enabling active learning and safety-critical use. They report strong apprenticeship-learning results across gridworlds, Lunar Lander, the Highway Environment, and two Atari games, spanning discrete and higher-dimensional settings. A practitioner could use the variational-Q posterior for uncertainty-driven demonstration querying or risk-aware policy extraction where standard maximum-entropy IRL gives no confidence estimate.

arXiv · cs.CVBuildable★ flagship

An Empirical Study of Training Pixel-Space Text-to-Image Diffusion Models

A recipe for making image generators paint directly in pixels, not a compressed shorthand.

Most top text-to-image AIs don't work on raw pixels — they generate in a compressed 'latent' space and then decode to a picture, which is faster but adds a lossy middleman. This paper hunts for a practical recipe to train models that generate directly in pixel space at large scale and still match or beat the compressed approach. Their key finding is that training in pixels from scratch is painfully slow to converge, so they use a two-phase trick: learn the general 'how images work' priors cheaply in latent space first, then transition to pixel space during later fine-tuning. They then carefully test the knobs that make that handoff work — how to initialize weights, what data to mix in, what the model predicts, the decoder design, and the noise schedule. It matters because a good pixel-space recipe removes the decoder bottleneck and could yield cleaner, higher-fidelity generation.

Technical view

The paper is an empirical study establishing a latent-to-pixel training recipe for large-scale text-to-image diffusion. It documents that direct large-scale pixel-space pre-training converges substantially slower than latent-space, motivating acquiring generative priors in latent space then transitioning to pixels in post-training. It systematically ablates the transition's design choices — weight initialization, data composition, prediction target (e.g. eps vs v vs x0), decoder architecture, and noise schedule — to make pixel-space models rival or exceed latent counterparts. Practitioners get an actionable set of transition hyperparameters and a staged-training strategy to drop the VAE decoder without paying the full from-scratch pixel convergence cost.

arXiv · cs.DSConceptual

Improving the matrix multiplication exponent with modern optimization and AlphaEvolve

AI just nudged the theoretical speed limit for multiplying giant matrices a tiny bit lower.

Multiplying two big grids of numbers (matrices) is one of the most basic operations in computing, and mathematicians have long hunted for the fastest possible way to do it, expressed as an exponent called omega — the smaller it is, the faster multiplication could theoretically go. This paper improves the best-known upper bound on that exponent, pushing it from 2.371339 down to 2.371177, by reworking a tricky internal optimization problem that underlies the leading method (the 'laser method' with combination loss analysis). They do this by reformulating the math to search a bigger space of solutions, inventing a new machine-learning-based search algorithm, and then polishing the result using AlphaEvolve, Google DeepMind's AI system for discovering better algorithms and proofs. It matters because even tiny improvements to this exponent ripple through theoretical computer science, since matrix multiplication underlies huge swaths of algorithms.

Technical view

The result targets the optimization core of combination loss analysis, the current state-of-the-art refinement of the laser method for bounding the matrix multiplication exponent ω (building on Duan et al. 2022, Williams et al. 2024, Alman et al. 2025). The authors reformulate the underlying optimization to be tractable over a larger search space than prior work, design a new ML-based optimization algorithm for it, and further refine solutions using AlphaEvolve, yielding ω < 2.371177 versus the prior best of 2.371339. Practitioners interested in fast linear algebra or algorithm discovery could study the reformulated optimization problem and the AlphaEvolve refinement loop as a template for attacking similar combinatorial/continuous optimization problems in theoretical CS.

arXiv · cs.DSConceptual

Spectral Gaps of Hit-and-Run and Coordinate Hit-and-Run

A random-walk trick for exploring shapes now provably mixes almost twice as fast as before.

Imagine you're trapped inside some weirdly-shaped high-dimensional room and want to bounce around randomly until your position is 'truly random' and uniform — this is exactly what algorithms need to do to sample from complicated probability distributions, which shows up in statistics, optimization, and machine learning. 'Hit-and-Run' is a classic random-walk method for this: at each step you pick a random direction and jump to a random point along the line in that direction within the shape. This paper proves a sharper mathematical bound on how quickly Hit-and-Run 'forgets' its starting point and converges to true randomness, showing the convergence speed depends on the shape's intrinsic properties (a quantity called the Poincaré constant) rather than a cruder measure of its size (the outer radius), and for well-behaved shapes this cuts the dependence on dimension from cubic to nearly quadratic. It matters because faster, better-understood sampling algorithms make simulations and statistical computations in high dimensions cheaper and more reliable.

Technical view

The paper establishes a spectral gap lower bound of Ω(1/(n²·C_PI)) for Hit-and-Run on any convex body containing a unit ball, where C_PI is the Poincaré constant of the target uniform distribution, implying O(n²·C_PI·log(M/ε)) mixing steps to reach χ²-divergence ε from arbitrary M-far starting distributions — refining the classical O(n²R²log(M/ε)) bound of Lovász and Vempala (2004) that depended on the body's outer radius R. Combined with progress on the KLS (Kannan-Lovász-Simonovits) conjecture, for nearly isotropic bodies this yields O(n²log n·log(M/ε)) complexity, improving dimension dependence from cubic to near-quadratic while preserving logarithmic dependence on initial distance; a companion result covers Coordinate Hit-and-Run. This is directly usable by anyone implementing MCMC samplers for convex-body sampling or log-concave distributions where tighter mixing-time guarantees translate to fewer iterations needed in practice.

arXiv · cs.SCBuildable

AutoSR: Automatic Symbolic Regression by Searching Research States

An AI system that finds equations from data now also keeps a lab notebook explaining its reasoning.

Symbolic regression is the task of finding a mathematical formula that fits a set of data points — useful in science when you want to discover the underlying law behind measurements. The problem is that with limited, noisy data, many different formulas can fit equally well numerically but behave wildly differently outside the range you measured, so just picking the 'best-fitting, simplest' equation isn't scientifically trustworthy. AutoSR tackles this by not just searching for equations but keeping a full 'Research State' for each candidate — a record of why it was tried, what evidence supports it, and independent critique — much like a scientist's lab notebook, maintained by AI agents that propose ideas and other AI agents that review them. This matters because it aims to make automated scientific discovery more like real science: traceable, justified, and self-correcting, not just curve-fitting.

Technical view

AutoSR instantiates 'Research-Space Symbolic Regression,' reframing the search from isolated candidate equations to persistent 'Research States' that bundle each equation with its motivating reasoning, computational evidence/probes, and independent review, generated through a proposer-reviewer multi-agent loop. This addresses a known failure mode where numerically competitive fits diverge wildly in extrapolation, since fit quality and syntactic complexity alone are insufficient credibility signals. By preserving the investigation history rather than discarding it after scoring, the system can inform later search branches with prior reasoning — a design pattern practitioners building agentic scientific-discovery or automated-hypothesis pipelines could adopt for maintaining provenance and justification alongside candidate outputs.

arXiv · cs.LGBuildable

An Analytical-Prior Framework for Data-Efficient Prediction of Sound-Reduction Frequencies in Rectangular Side-Branch Helmholtz Resonators

Blending old-school physics formulas with modern AI lets you predict noise-cancelling resonators from fewer simulations.

Helmholtz resonators are cavity-and-neck devices (like blowing across a bottle top) used to cancel out unwanted sound at specific frequencies in ducts and pipes, and engineers want to predict exactly which frequencies a given design will block. Getting precise predictions normally requires expensive, slow computer simulations (finite-element analysis), and while you could train a machine-learning model on simulation data instead, that model becomes unreliable if you don't have many simulations to learn from. This paper's fix is to lean on an existing, cheap, rougher analytical formula for these resonators as a 'prior' — a starting guess — and let the limited simulation data teach the ML model only the gap between that rough formula and reality, rather than learning the whole physics from scratch. This matters for engineering practice because it means accurate noise-control predictions can be made even when you can't afford to run thousands of expensive simulations.

Technical view

The paper proposes an analytical-prior learning framework for predicting sound-reduction (transmission loss) frequencies of rectangular side-branch Helmholtz resonators under limited high-fidelity FE simulation budgets. Two variants are given: a discrepancy-learning route that keeps the analytical model as an explicit baseline at inference and trains the learner only on the analytical-to-simulation residual, and a distillation route that pretrains a surrogate on abundant cheap analytical evaluations before calibrating it with scarce simulation data for cases needing a self-contained predictor. This hybrid physics-informed/data-efficient approach is directly applicable to other acoustic or structural design problems where a low-fidelity analytical model exists alongside expensive high-fidelity simulations, offering a template for reducing simulation-data requirements via residual or distillation-based transfer.

arXiv · cs.LGBuildable

Data-Efficient and Interpretable Classification of Circulating Tumor Cell Phenotypes in Microfluidic Devices via Deep Learning

A tiny chip sorts cancer cells by how they tumble through fluid, and AI reads the tumbles.

Circulating tumor cells (CTCs) are cancer cells that break off a tumor and travel in the blood, and their shape and squishiness hint at how dangerous they are. Researchers push cells through a maze-like microfluidic chip — basically a tiny obstacle course carved into a chip — and each cell's size and stiffness bends its path in a distinctive way. Untangling which path shape came from which cell type is really hard to do with plain physics equations, so the team built a deep learning model to learn the pattern instead, but designed it to need only a little training data and to explain its own reasoning rather than being a black box. This matters because doctors could eventually use a cheap chip plus this AI to quickly read a patient's cancer risk from a blood draw instead of invasive biopsies.

Technical view

The work targets the inverse problem of recovering CTC biophysical phenotype (size, deformability) from trajectory data generated by label-free microfluidic sorting, where fluid-structure interaction nonlinearity makes analytical inversion intractable. The authors propose a data-efficient DNN architecture with built-in interpretability constraints, addressing the typical trajectory-data scarcity that limits standard black-box classifiers. This suggests physics-informed or feature-constrained network design tying kinematic trajectory features back to physically meaningful cell properties, which practitioners could extend to other microfluidic sorting/classification tasks facing similar low-data regimes.

arXiv · cs.CLConceptual

Towards Computational Provenance: Carrying Causal-State Evidence in Generated Text

Can AI text secretly prove which internal 'thought path' actually wrote it?

When a language model answers a question, you have no way to verify what actually happened inside its 'brain' to produce that answer — you just get the output. This paper explores whether the model could leave a hidden, checkable trail of evidence about its own internal reasoning path baked right into the text it writes. The researchers tested this on simple architectures doing arithmetic where the model could take one of two valid internal routes to the same correct answer, deliberately forcing it to switch between routes and checking whether a faint statistical fingerprint in the generated text reliably revealed which route was actually used. It worked perfectly across all their test cases, which matters because it's a first step toward trusting AI systems by verifying their internal decision-making, not just their final output.

Technical view

The authors define 'computational provenance' — embedding detectable evidence of which causal internal state produced a model's output directly into the generated text — and test it on a controlled arithmetic task with two mutually exclusive discrete intermediate computation paths, in both a modular feed-forward network and a transformer. By deliberately switching the active internal path and authenticating it against a subtle statistical signature in the output text, both architectures passed all 128 matched public/held-out verification pairs. This is essentially a steganographic self-attestation mechanism for model internals, and could be built on for auditing or detecting which reasoning circuit/path a model actually used at inference time, relevant to interpretability and AI safety verification work.

arXiv · stat.MLBuildable

Non-Crossing Deep Quantile Regression for Distributional Survival Prediction

A survival-prediction AI that promises its risk curves will never contradict themselves.

In medical and reliability studies, researchers often want to predict not just 'when will this event happen on average' but the whole range of possible timings — like the fastest 10% of cases versus the slowest 10%. The problem is that many patient factors affect early risk differently than late risk, and standard methods squash all that nuance into one number, while flexible methods that try to capture multiple time points can produce nonsensical results where a 'later' prediction accidentally comes out earlier than an 'earlier' one. This paper builds a system that estimates several of these time-point predictions at once while mathematically guaranteeing they stay in the correct order, using flexible modern neural network components (including a newer type called Kolmogorov-Arnold networks, plus Transformers) so it can still capture complex patterns. It matters because more reliable, self-consistent risk timelines could improve how doctors and engineers plan for events like disease progression or equipment failure, especially when some outcomes are never observed (censored) in the data.

Technical view

The paper introduces a Censored Non-crossing Quantile (CNQ) framework for right-censored survival data that jointly estimates multiple conditional quantiles of event time with a hard ordering constraint built into the model architecture, avoiding the crossing-quantile pathology of flexible quantile regression. Flexibility comes from Kolmogorov-Arnold Network and Transformer backbones, and the authors provide a finite-sample excess-risk bound holding jointly across all fitted quantile levels — a theoretical guarantee uncommon in deep survival modeling. Evaluated on 27 simulation settings plus six real cohorts, CNQ beats quantile-, hazard-, and mean-based baselines on pinball loss, making it a candidate drop-in for practitioners needing calibrated, monotonic conditional survival distributions rather than single point estimates.

arXiv · cs.CVBuildable

SplatGuide: Geometric Priors from 3D Gaussians for Pose-Free Novel View Synthesis

Reusing one 3D scene reconstruction three different ways to fill in never-seen camera angles.

Imagine taking a handful of photos of a room from random, unknown angles and wanting a computer to generate a photorealistic view from a totally new angle you never photographed. Current systems typically build a rough 3D model of the scene using something called 3D Gaussian Splatting (like sculpting the scene out of thousands of tiny colored blobs) and then use an image-generating AI to fill in gaps, but they usually only pull one piece of information out of that 3D model — either just the rendered picture or just internal features — wasting other useful information like knowing which parts of the scene are hidden behind objects. SplatGuide instead squeezes three complementary signals out of the same 3D reconstruction: a rendered guide image, a map showing which original photo best shows an occluded area, and internal feature representations, feeding all three into the generation process together. This more efficient reuse of the same 3D data should produce sharper, more consistent new views without needing to know precise camera positions ahead of time, useful for things like virtual walkthroughs or AR from casual photo sets.

Technical view

SplatGuide targets pose-free novel view synthesis by combining feed-forward 3D Gaussian Splatting (3DGS) reconstruction with multi-view diffusion, addressing the limitation that prior pipelines extract only a single signal (either rendered pixels or learned features) from the 3DGS reconstruction. The method reuses one 3DGS scene in three roles simultaneously: pixel-aligned geometric conditioning from rendered images, an occlusion-aware target-view voting map built from per-Gaussian source-view visibility indices for reference selection, and feature-level guidance via cross-attention from reconstruction tokens. This multi-signal fusion is positioned as closing an 'information disconnect' in existing feed-forward-3DGS-plus-diffusion pipelines, giving practitioners a template for extracting richer conditioning signals from a single Gaussian splat reconstruction rather than discarding visibility and feature information after rendering.

arXiv · cs.DMConceptual

The canonical facets of multi-separator polytopes

A math toolkit pins down the sharpest rules for slicing images into clean regions.

Splitting a photo into meaningful regions — like separating a tumor from healthy tissue — can be framed as a puzzle over a graph, where 'cuts' separate different areas. This paper studies the exact shape of the solution space for a specific version of that puzzle called the multi-separator problem, a newer alternative to an older 'lifted multicut' approach. The researchers work out precisely which mathematical inequalities form the true boundaries ('facets') of that solution space, using conditions a computer can check efficiently. Nailing down these boundaries matters because it lets optimization software solve segmentation problems faster with provable guarantees instead of rough approximations.

Technical view

The paper gives a polyhedral analysis of the multi-separator ILP (Irmai et al., 2024), characterizing which classes of constraints define facets of the multi-separator polytope via efficiently-decidable graph-theoretic conditions. It strengthens these inequalities to yield additional facets, and for path graphs with all-pairs separation obtains a totally dual integral (TDI) description — a strong integrality/optimality guarantee. It also connects the polytope to the boolean quadric polytope, showing odd-cycle inequalities carry over as facets. Practitioners building branch-and-cut solvers for segmentation can use these facet characterizations directly as cutting planes to tighten LP relaxations.

arXiv · cs.CVBuildable

HarnessEval-W: Agentifying the Evaluation of Visual Worlds

An AI judge that shows its work when grading whether simulated worlds obey physics.

World models are AI systems that simulate video-game-like environments, predicting how a scene will evolve — but until now there was no good way to check if those predictions make physical sense, just an opaque single score. HarnessEval-W fixes this by acting like a detective: it breaks a big evaluation question, like 'did gravity behave correctly here?', into smaller sub-questions, then sends specialized AI helper agents, each with the right tools, to investigate one piece and report back with actual reasoning. This produces a chain of evidence a human can inspect and verify, rather than a black-box number. It matters because as world models get used to train robots and self-driving systems, we need to trust their evaluation, not just its score.

Technical view

HarnessEval-W adapts the 'harness' pattern from LLM agent evaluation to world-model benchmarking: instead of a fixed metric pipeline, it decomposes each evaluation case into measurable subproblems and dispatches specialized sub-agents, equipped with diagnostic tools, to independently assess physics, causality, and state-consistency violations in a rollout. Each sub-agent produces an inspectable reasoning trace that is aggregated into the final judgment, giving auditability that brute-force metric computation lacks. This is directly usable as a drop-in evaluation harness for video/world-model generation research, replacing fixed rubrics with agentic, tool-using judges. Builders could extend it by adding new diagnostic tools or sub-agents for domain-specific physical rules.

arXiv · q-fin.RMBuildable

zLend: A Dual-Scope Cash-Flow Reconstruction Framework for On-Chain Credit Underwriting

Reading crypto wallets like bank statements to score who's creditworthy without a credit bureau.

If you want to lend crypto to a stranger, there's no credit bureau to check whether they'll pay you back — only their public blockchain transaction history. zLend reconstructs a wallet's day-by-day cash balance from raw transfer records, twice over: once counting only stable, dollar-pegged tokens, and once counting all their crypto assets, because holding lots of volatile tokens isn't the same as having spendable cash. From these two views it computes signals like whether the wallet has enough liquid money to cover a loan, how steady its finances are, how well it bounces back from big financial dips, and whether it has recurring trusted counterparties, similar to a regular paycheck. It's a real, deployed system bringing something like traditional credit-risk analysis into decentralized finance.

Technical view

zLend reconstructs a per-wallet daily balance time series from on-chain token transfer logs, computed over two distinct scopes (a fixed stablecoin basket vs. all fungible transfers) to separate liquid spendable balance from total token holdings. From each reconstructed series it derives underwriting features: liquidity coverage against a target loan size, cash-flow volatility and regularity, a drawdown-and-recovery statistic adapted from quantitative finance, and a recurring-counterparty detector acting as a proxy for income regularity. It's presented as a deployed production framework, so DeFi lending protocols could adopt similar dual-scope reconstruction as a risk-scoring input, with the drawdown/recovery and volatility features portable to other on-chain risk models.

arXiv · cs.CVBuildable

Can Unsupervised Methods Outperform Supervised Deep Learning When Ground Truth Is Sparse? A Case Study of Bronchovascular Bundle Segmentation in Low-Dose CT

Can old-school unsupervised math beat deep learning at mapping lungs when labels are scarce?

Lung cancer is often caught too late, and even when a suspicious spot exists, tiny tumor nodules can hide behind the tangle of blood vessels and airway walls right next to them. Deep learning usually needs lots of hand-labeled scans to learn to trace these structures, but radiologists don't have time to label enough low-dose CT images by hand. This study asks whether older 'unsupervised' methods — which don't need labeled examples, just clever rules about the image data — can match or beat deep learning at outlining these vessel-and-airway bundles. They test both approaches on real lung-cancer screening datasets to see which one actually clears the way for spotting nodules faster. It matters directly for easing the growing backlog of scans radiologists face.

Technical view

The paper compares unsupervised segmentation methods against supervised deep learning for delineating the bronchovascular bundle (airway walls plus adjacent vessels) in low-dose CT lung screening, motivated by scarce ground-truth labels in this domain. It evaluates on established LDCT screening cohorts, including the Duke Lung Cancer Screening dataset and the Pilot Pomeranian Lung Cancer Screening Program, using bundle isolation/removal as a preprocessing step meant to improve downstream nodule detectability. The core contribution is an empirical case study rather than a new architecture, directly testing whether classical unsupervised segmentation can match supervised deep nets under label scarcity. CAD pipeline developers could adopt whichever approach wins as a labeling-free preprocessing stage ahead of nodule-detection models.

arXiv · cs.AIConceptual

What Do Compliance Detectors Read? An Audit of Activation Probes and Guard Models

AI compliance checkers often ignore the actual rule and just react to how risky a scenario feels.

Companies increasingly use AI 'guard' models to check whether another AI's output breaks a written rule, like a privacy law or hospital policy, and treat that check as a real audit control. This paper shows that trust is often misplaced: if you delete the rule entirely, swap it for its opposite, or scramble its wording, these detectors give almost the same verdict anyway, meaning they aren't actually reading the rule at all. Even a fancier guard that correctly quotes the right legal clause in its explanation barely changes its decision when that clause is flipped to mean the opposite. They prove this with a clever test that crosses two different rules with two different scenarios, so no detector can cheat by just recognizing 'this looks risky' instead of genuinely applying the rule. It matters because it suggests today's AI compliance tools may be rubber-stamping decisions rather than truly enforcing policy.

Technical view

The authors define 'rule blindness': a compliance detector's verdict should depend on the stated governing rule rather than scenario surface features, but this fails across activation probes and guard-model detectors, including a policy-conditioned guard that cites the correct clause yet barely shifts its verdict when that clause is swapped for its logical opposite. Their diagnostic uses a 2x2 benchmark crossing two rules with two scenarios so neither the rule nor the scenario alone predicts the correct label, isolating true rule-sensitivity from confounds. Ablations (deleting, permuting, or substituting the rule text) leave detection accuracy essentially unchanged, indicating detectors rely on spurious scenario correlates rather than rule semantics. This gives auditors a concrete stress-test methodology to apply to any deployed compliance guard or probe before trusting it as an audit control.

arXiv · cs.LGBuildable

Proteus: Incremental Memory Activation for Long-Context Sequence Modeling

A long-context AI memory that starts tiny and grows, so early clutter doesn't crowd out later info.

AI models reading very long documents or conversations often compress everything into a fixed-size memory, but if that memory has full capacity from the very first word, early details grab too much space and 'pollute' it, leaving little room for what matters later. Proteus instead starts the memory deliberately cramped, forcing the model to ruthlessly compress the beginning, then gradually opens up more capacity as more context arrives. It's a bit like a student jotting only the gist of chapter one but taking fuller notes later once they know what's actually important. The result is a model that retains recent, relevant information better and suffers less interference between old and new content, which matters for anything processing huge documents, codebases, or long conversations efficiently.

Technical view

Proteus introduces 'incremental memory activation' for memory-based (recurrent/state-space style) long-context sequence models: rather than a static full-capacity memory throughout the sequence, effective memory capacity is progressively expanded as context grows, imposing an early compression bottleneck and unlocking capacity later. This reduces early-token 'pollution' — over-allocation of memory degrees of freedom to early tokens — and lowers interference between stored history and incoming tokens, improving retention of later context. It's framed as a general paradigm instantiated in a concrete architecture, positioned against static-memory baselines in the memory-compression literature that seeks alternatives to quadratic attention. Practitioners building long-context memory architectures could apply the same progressive-capacity schedule to their own recurrent memory modules as a training curriculum change.

arXiv · cs.ROBuildable

HAF: Adapting Generalist VLAs to Humanoid Whole-Body Loco-manipulation via Hierarchical Action Flow and Spectral Latent RL

Teaching general-purpose robot brains to walk and use both arms together, safely.

Generalist 'vision-language-action' AI models can tell a robot arm what to do from a photo and a text instruction, but humanoid robots are much harder: they must coordinate walking, balancing their waist, and using two arms all at once, which single-stage AI models struggle to manage. Robots trained purely by imitating human demonstrations often perform suboptimally once deployed, and letting them learn by trial-and-error directly on a giant AI brain is expensive and risky, since it might fall over while exploring. HAF splits the problem in two: one part turns a generalist model's plan into a layered sequence of actions that coordinates the whole body, and a second, lighter part fine-tunes behavior efficiently through trial-and-error without retraining the entire giant model. This matters for getting humanoid robots to reliably do useful physical tasks in homes and workplaces.

Technical view

HAF (Humanoid Adaptation Framework) is a two-part system adapting generalist VLA foundation models to humanoid whole-body loco-manipulation: HAF-VLA produces a hierarchical action flow that decomposes and coordinates locomotion, waist posture, and dual-arm manipulation from a single VLA plan, addressing the coordination failure of single-stage VLA architectures on high-dimensional humanoid action spaces. A second component performs efficient online reinforcement learning in a compressed latent space (spectral latent RL) to refine the behavior-cloned policy after deployment, avoiding the cost and safety risk of directly fine-tuning the full VLA backbone. This targets two known failure modes — poor multi-DOF coordination in monolithic VLA policies and the suboptimality of offline behavior cloning — via a modular hierarchy-plus-latent-RL approach. Robotics practitioners adapting VLA models to new embodiments could reuse the hierarchical action-flow decomposition and latent-space RL fine-tuning stage as a template for other high-DOF robots.

arXiv · cs.CLConceptual

Model Hypnosis: Strong control of AI via additive subliminal effects

Tiny, innocent-looking typos and rewordings can secretly gang up to hijack an AI's behavior.

You'd think a few random typos or slightly reworded sentences in a prompt couldn't meaningfully change what an AI does, since each one alone seems totally harmless. This paper shows that's wrong: many individually weak, inconspicuous cues, like specific paraphrasing choices or small typos, can be combined to strongly and reliably steer an AI model's behavior, a phenomenon the authors call 'model hypnosis.' It works across different AI companies' models and sizes, including the most advanced reasoning models, and prompts crafted to hypnotize one model often work on totally different models too. This is concerning because someone could covertly manipulate an AI's output using text that looks completely innocent to a human reader or another AI checking for obviously suspicious prompts, undermining both safety filters and our ability to understand why a model behaved as it did.

Technical view

The paper demonstrates 'model hypnosis': additive combinations of individually weak, semantically-irrelevant textual cues (paraphrase choices, typos, stylistic variations) compose into prompts that exert strong, reliable control over model outputs. The effect generalizes across model families and scales, holds for frontier reasoning models, and hypnotic prompts show cross-model transferability, suggesting the underlying subliminal-signal exploitation isn't model-specific overfitting. Because the controlling signal is distributed across many innocuous surface features rather than concentrated in an obviously adversarial token sequence, standard prompt-injection defenses and interpretability tools that look for salient triggers are likely to miss it. This gives red-teamers a concrete technique to probe deployed models for susceptibility to additive subliminal cues, and hands interpretability researchers a new phenomenon — distributed subliminal control — to try to detect or explain mechanistically.

arXiv · cs.LGBuildable

Time-Aware Validation of Machine Learning Fuel Consumption Models: Evidence from 1\,Hz Operational Data, CCGS \textit{Sir Wilfrid Laurier}

How you test a ship-fuel AI model secretly decides whether it actually works.

Ships burn fuel constantly, and companies want AI models that predict how much fuel a vessel will use so they can plan routes and cut emissions. But there's a hidden trap in how these models get graded: most researchers test them by randomly shuffling their data into 'training' and 'testing' piles, which accidentally lets the model peek at future information it shouldn't have, like a student seeing tomorrow's exam answers today. This paper instead tests fuel-prediction models the honest way, only ever letting them learn from the past to predict the future, using real minute-by-minute data from a Canadian Coast Guard icebreaker. The result matters because a model that looks great under the old sloppy testing might actually flop once it's deployed on a real ship.

Technical view

The study benchmarks six regression models plus a physics-based baseline for ship fuel consumption (SFC) prediction using 1Hz operational data from CCGS Sir Wilfrid Laurier, comparing standard random-split validation against time-aware schemes: Time Series Cross-Validation (TSCV) and Blocked TSCV (BTSCV). Models are tuned across three time-aware validation schemes and three feature configurations to quantify how much random splitting inflates performance via temporal leakage. Practitioners building maritime DSS or emissions-estimation tools can use the paper's protocol to avoid overstating deployment-time accuracy, and the blocked/time-series CV setup is directly reusable for any high-frequency time series sensor modeling task.

arXiv · cs.AIConceptual

Policy Iteration with Human Feedback: Bringing Post-Training RL to In-context Learning

An AI keeps a living rulebook of its own mistakes, edited by human experts, not retrained from scratch.

Normally, improving an AI system after it's deployed means expensively retraining the whole model. This paper proposes a lighter alternative: instead of changing the AI's brain, you keep a written 'policy' document, essentially a set of instructions and tools, that the AI follows, and you revise that document over time as problems crop up. An AI critic and a human clinical expert review the system's reasoning and actions, spot recurring mistakes, and propose edits to the policy, but the human expert always has the final say on whether a change sticks or gets rolled back. This matters because it offers a cheaper, more controllable, and more transparent way to keep AI systems improving in high-stakes fields like medicine, where you want a human checking every update.

Technical view

Policy Iteration with Human Feedback (PIHF) reframes post-training RL's generalized-policy-iteration loop as edits to a versioned natural-language policy and toolset rather than gradient updates to model weights, using a frozen pretrained LM as the execution substrate. An LM critic plus a clinical expert jointly review complete reasoning/tool-use trajectories to localize recurrent failure modes and draft candidate policy revisions, with the human expert retaining authority over admission and rollback; Recall@1/Recall@5 are used as outcome metrics after candidate execution, evaluated via cumulative ablations on ultra-rare-disease diagnosis tasks. This is essentially in-context/prompt-level policy improvement with a human-in-the-loop gatekeeper, offering a reusable pattern for auditable, reversible agent improvement without weight updates.

arXiv · cs.LGRunnable

CaliBench: Are the Stochastic Dynamics of Video World Models Physically Calibrated?

A test that checks if AI-generated videos roll dice and flip coins with the right odds.

Video-generating AI models are increasingly used to simulate physics, like predicting how a ball might bounce or fall, but existing tests only check if a single video looks realistic or roughly matches a dataset on average. This paper builds a benchmark called CaliBench that instead checks something more precise: if you ask the AI to generate many outcomes of a truly random physical event, like a Galton board where balls bounce into bins in a bell-curve pattern, or a spin of a roulette wheel, do the odds it produces actually match the real known probabilities? They use situations where the correct answer is mathematically known in advance, like dice or cards, so they can grade the AI's 'randomness' exactly rather than just eyeballing realism. This matters because if AI world models are going to be used for planning or simulation, they need to get not just plausible-looking outcomes but the correct statistical spread of outcomes.

Technical view

CaliBench evaluates video world models' stochastic sampling by scoring generations in interpretable discrete outcome spaces (bin index, die face, card suit, roulette color) with closed-form reference distributions (binomial Galton boards, Bernoulli forks, uniform dice/cards/lottery, skewed roulette), rather than relying on feature-space metrics like FID that only coarsely compare distributions. The framework decomposes performance into two orthogonal axes — scorability (fraction of generations that yield a determinable discrete outcome) and calibration accuracy against the known reference distribution — separating 'can we even read an answer out' from 'is the answer's probability correct.' This gives builders of video world models a precise, ground-truth-backed diagnostic for aleatoric uncertainty calibration, usable as a regression test during model development.

arXiv · cs.LGBuildable

GEO-Flag: Detecting and Measuring GEO-Optimized Web Content

A detector for web pages secretly written to trick AI search engines into citing them.

When you ask an AI search engine a question, it reads webpages and synthesizes one direct answer instead of showing you a list of links to check yourself. That creates an incentive for website owners to write content specifically designed to get quoted by the AI, called Generative Engine Optimization or GEO, similar to how people once wrote pages to rank higher on Google. The danger is that a page optimized this way could get cited by the AI even if it's low-quality or outright wrong, because the AI never really vets its sources the way a human clicking through links might. This paper builds a large test set of thousands of webpages, some GEO-optimized and some not, across many topics and optimization techniques, and uses it to check how well current detection tools can catch this manipulation. It matters because AI answers are only as trustworthy as the pages feeding them, and right now there's little defense against gaming that pipeline.

Technical view

GEOFlagBench is a benchmark of 3,200 webpages across 400 queries, four domains, and eight distinct GEO-optimizer families, built to systematically stress-test methods for detecting generative-engine-optimized content. The authors evaluate existing GEO detection baselines against this benchmark and find performance gaps (the abstract notes even the strongest baseline underperforms, cut off before specifics), establishing a reproducible testbed for the emerging adversarial problem of provenance and authority laundering in AI-synthesized search answers. Researchers building content-provenance or search-integrity tooling can use this benchmark directly to train or evaluate classifiers distinguishing organically authoritative content from strategically optimized text.

arXiv · cs.CVBuildable

Unlocking the Potential of Image Editing via Concept Scaling and Dense Supervision

A 12-million-example dataset teaches AI to edit fine details in photos without messing up the rest.

AI image editors, ones that let you type 'change the sky to sunset' and have it happen, are usually built the same way as AI image generators, but editing is actually a different problem, requiring precision on a specific small change rather than creating a whole scene from scratch. This paper points out two weaknesses in how these editors are trained: they don't distinguish between coarse edits, like changing a whole scene, and fine ones, like adjusting a single button on a shirt, and each training example usually only teaches the model one change at a time, wasting a lot of potential learning signal. Their fix is to build a giant catalog of over 1,000 specific types of edits and 12 million example before-and-after image pairs, then train with images that pack multiple non-conflicting edits into a single example so the model learns more from each picture. This means future AI photo editors could get noticeably better at precise, complex, multi-part edits rather than vague or clumsy ones.

Technical view

The authors identify two failure modes in adapting text-to-image diffusion training to image editing: coarse concept granularity and sparse per-sample supervision. Their solution combines a hierarchical taxonomy of 1,000+ fine-grained edit concepts with ConceptEdit-12M, a 12M-pair dataset generated via an improved library-driven synthesis framework that mitigates distribution collapse common in synthetic training data, plus a dense supervision strategy that composites multiple non-interfering edit concepts into single training pairs to increase signal density per sample. This is a data-and-training-strategy contribution rather than a new architecture, so practitioners fine-tuning diffusion-based editors could adopt the taxonomy and dense-supervision composition technique directly to improve edit fidelity and multi-concept editing without architectural changes.

arXiv · cs.ROConceptual

When State Becomes an Attack Surface: State-Semantic Injection in LLM-Driven Embodied Agents

Hackers could poison a robot's own memory of the world to make it act against its owner.

Robots controlled by AI language models don't just follow commands blindly, they build an internal picture of their surroundings, called 'state', tracking things like where objects are and what's already been done, and use that picture to decide what to do next. This paper points out that this internal state itself can be attacked: if someone can slip false information into what the robot believes about its environment, they can manipulate its next actions without ever touching its code or sensors directly, similar to gaslighting a person by lying about what's in the room. The researchers study this as a new kind of security weakness specific to AI-driven physical robots, distinct from typical hacking that targets software or networks. It matters because as robots increasingly rely on language-model 'brains' to plan and act in the real world, an attacker manipulating their beliefs, rather than their hardware, becomes a genuinely new and hard-to-detect threat.

Technical view

The paper defines State-Semantic Injection, an attack surface in LLM-driven embodied agents where an adversary corrupts the natural-language or symbolic 'state' representation (the agent's running belief about environment, task progress, and affordances) that mediates between perception and action, rather than attacking perception, tools, or the LLM's weights directly. It situates this within the SayCan/Code-as-Policies/ProgPrompt lineage where LLMs perform task decomposition and planning over structured or semantic state, arguing that because these systems trust their internal state representation as ground truth, injecting falsified semantic state can hijack downstream planning and action generation. This establishes a threat model and likely an attack methodology/benchmark (per the abstract's framing) that security researchers evaluating embodied-agent robustness could use to design state-integrity defenses or detection mechanisms.

arXiv · cs.CVBuildable

Diagnosing Dense Same-Class Attribute Misbinding in Large Vision-Language Models

AI vision models can spot a red shirt and a person, but still put the red on the wrong person.

When there are several similar objects in a photo, like a group of people or a row of cars, AI vision-and-language models can correctly notice both the objects and their features, say, that someone is wearing blue and someone else red, but then mix them up and describe the wrong person as wearing blue. Existing tests don't catch this specific mistake: a general question-answering test just marks the answer wrong without saying why, and a hallucination test might even judge the answer fine since both the object and the attribute genuinely appear somewhere in the image, just attached to the wrong instance. This paper names this exact failure 'Dense Same-Class Attribute Misbinding' and builds a large, carefully controlled test set of crowded scenes with labeled individual objects and their true attributes to directly measure how often and why this crossed-wires mistake happens. This matters because these mix-ups are a subtle but important limitation for any application, like surveillance or assistive tech, that depends on correctly matching descriptions to the right individual in a crowd.

Technical view

The paper formalizes Dense Same-Class Attribute Misbinding (DSCAM), where VLMs correctly perceive object and attribute presence but bind an attribute to the wrong same-class instance, and introduces InstaBind-Lite: 524 images with 529 curated groups of 3-6 same-class entities, 1,773 boxed instances, ordered spatial neighbors, and distinguishable color-like attributes, yielding 9,580 deterministically evaluated questions across four question complexity levels. Critically, the benchmark uses source-instance annotations to disentangle three distinct error types — unsupported generation, plain recognition failure, and true attribute misbinding (attribute correctly recognized but copied from a neighboring instance) — which standard VQA accuracy and object-hallucination metrics conflate. This gives VLM developers a diagnostic tool to isolate and specifically target binding/grounding failures in dense scenes, separate from general hallucination or recognition fixes.

arXiv · cs.AIBuildable

Cross-Sign Language Transfer Learning Using Domain Adaptation with Multi-scale Temporal Alignment

Teaching a computer American Sign Language by having it learn from other, unrelated sign languages first.

There are over 100 different sign languages worldwide, but nearly all the good training data and AI resources exist for just a handful, leaving most sign languages with almost no AI recognition tools. This research tackles that gap using 'transfer learning', taking what an AI already learned from data-rich sign languages and adapting it to a target language with little data of its own, specifically using a technique called domain adaptation, which nudges the AI's understanding to shift from one language's patterns to another's rather than starting from zero. They found that a specific domain-adaptation method that aligns how signs unfold over time at multiple time-scales worked better than plain transfer learning, especially for improving American Sign Language recognition, and that matching shorter clips of motion between the source and target worked best. They also compared using plain video color versus 'optical flow' (motion-only data) and found regular color video actually worked better overall. This kind of research matters because it could make sign language recognition technology accessible for the many sign languages that currently have almost no AI support.

Technical view

The work applies TA3N (Temporal Attentive Adversarial Adaptation Network) domain adaptation, which leverages a Temporal Relational Network (TRN) module to align multi-scale temporal relations, for cross-sign-language transfer, comparing it against standard neural transfer learning baselines for improving American Sign Language (ASL) recognition from other sign language source domains. Key findings are that domain adaptation outperforms plain transfer learning, that aligning shorter-term temporal relational features between source and target domains is most effective, and that RGB input modality outperforms optical flow across most experimental conditions. Practitioners building low-resource sign language recognition systems can reuse the TA3N/TRN pipeline and RGB-preference finding as a starting configuration when data for a target sign language is scarce but a related source-language dataset is available.

arXiv · cs.CLBuildable

ClawGym II: Exploring Black-Box RL on Agent Harness

Training AI agents to get smarter by learning through the messy software scaffolding that runs them.

An 'agent harness' is the scaffolding that lets an AI use tools, browse the web, or run code across many steps to finish a task. Training these agents with reinforcement learning (reward-based trial and error) is hard because the harness is a long, opaque, multi-step black box that's expensive to run at scale. The researchers built isolated sandboxes so thousands of task attempts can run simultaneously, and a proxy that quietly watches everything the AI says to the harness so it can reconstruct the full decision history for training. This matters because it's a general recipe for making AI agents better at long, real-world jobs without having to rewrite the harness software itself.

Technical view

The framework decouples RL policy optimization from opaque harness execution by placing a serving proxy at the model boundary to intercept all model calls, while a sandbox-based execution layer isolates task environments and harnesses for large-scale concurrent rollouts. Captured calls are reorganized into prefix trees to reconstruct multi-turn trajectories and boost training efficiency. This is directly applicable to RL fine-tuning of agentic LLMs operating through arbitrary third-party or in-house harnesses without needing white-box access to harness internals.

arXiv · cs.IRBuildable

UniDot: A Unified Network for Sequence Modeling and Feature Interaction in Large-scale Recommendation

One math trick lets a single model replace two separate designs powering online recommendations.

Recommendation systems (like the ones deciding what you'll click next) traditionally use two different kinds of models: one that mixes together your profile features like age or device, and another that studies your past behavior as a sequence of actions, loosely bolted together in production. The key insight here is that the core math operation in the older feature-based approach — multiplying two vectors together — is secretly the same operation as 'attention,' the mechanism behind modern sequence models. So the researchers built a single architecture, UniDot, that turns both feature and behavior data into a shared format of 'tokens' and processes them together in one stack. This matters because it could simplify and improve the recommendation engines that power e-commerce and content platforms.

Technical view

UniDot reframes the FM (factorization machine) embedding inner product and Transformer attention's query-key dot product as the same primitive, then tokenizes non-sequential fields and multi-domain behavioral sequences into one shared token space. A stacked macro-block combines a token-mixing bus (feature interaction) with a sequence-retrieval bus (cross-attention over item tokens) for post-click conversion prediction. Practitioners running dual-stack industrial recommenders could use this as a blueprint to merge feature-interaction and sequential-modeling pipelines into a single trainable network.

arXiv · cs.CEConceptual

Historical Backtesting for Scientific Question Discovery: A Protocol and Astronomy Pilot

Testing an AI's scientific guesses by rewinding the clock and checking who was actually proven right.

Some AI systems try to generate good scientific research questions, but until now there's been no objective way to check if a generated question is actually good — just subjective expert or AI ratings. This paper's fix: freeze the AI's knowledge at some past date, let it propose questions, lock those questions in before it can see anything published later, then compare them against the real literature that came out after that date to see whether each question got answered, partly addressed, independently asked by someone else, or ignored. It's like a scientific version of backtesting a stock-picking strategy. This matters because it turns 'is this a good research idea' into something falsifiable and reproducible, and they've released a working astronomy test kit for anyone to try it on.

Technical view

The protocol freezes a corpus at a historical cutoff, generates and locks candidate questions before exposure to later literature, then uses a temporally isolated future corpus to label each question's fate (answered / partially addressed / independently posed / ignored) and whether its premise was supported or refuted. It's model-agnostic — any system emitting frozen questions can be scored against it. The released astronomy instance includes frozen questions, auditable labels, four reference baselines, and a submission interface, so practitioners can plug in new question-generation systems and get comparable, falsifiable scores instead of subjective LLM-judge ratings.

arXiv · cs.ROBuildable

Neurosymbolic Embodied Agents

A robot that plans chores by combining AI hunches with airtight logical rules so the plan actually works.

When you ask a language or vision AI to plan a household task like 'make coffee,' it can produce a plausible-sounding plan that secretly breaks the rules of physics or refers to objects that aren't really there. This system splits the job in two: first, an AI camera-equipped explorer wanders the room and interacts with things to build an accurate, formal list of what objects exist and their properties (a 'symbolic' map of the world); second, a strict rule-based planner only considers actions that are actually valid given that map, searching through many possible move sequences (a technique called Monte Carlo tree search) to find one that's guaranteed to be executable. This matters because it fixes a core weakness of AI planners — plans that look right but fail in reality — by building in a formal safety net.

Technical view

Phase one uses a VLM plus an exploration harness to ground goal-relevant predicates and instance bindings from egocentric observations and interactions into a symbolic initial state. Phase two uses a PDDL transition model to restrict decoding to tokens extending applicable actions, with Monte Carlo tree search evaluating executable continuations via a domain-independent planning heuristic, producing plans that are executable by construction. Practitioners building embodied agents can adopt this pattern — VLM grounding feeding a constrained classical planner — to get formal executability guarantees instead of relying purely on end-to-end LLM plan generation.

arXiv · cs.CVBuildable

PixRestore: Unified Image Restoration via Pixel Diffusion Transformer

An AI photo-fixer that works directly on raw pixels instead of borrowing shortcuts from image generators.

Image restoration tools try to fix all kinds of damaged photos — blurry, noisy, low-resolution — with one model. Many recent approaches repurpose huge pretrained text-to-image generators, but those models compress images first (losing fine detail) and can add hallucinated content that doesn't match the original. PixRestore instead trains a new model completely from scratch that works directly on raw pixels in small patches, using a technique called flow matching to gradually transform noise into a clean image, which keeps computation manageable while preserving fine detail. It also learns to judge how trustworthy different internal signals are depending on the type of damage. This matters because it avoids the detail-loss and hallucination problems that come from reusing generic image-generation models for restoration.

Technical view

PixRestore is a VAE-free, pixel-space Diffusion Transformer trained from scratch (no T2I pretraining), performing flow matching directly on patchified pixels rather than a compressed latent space, which keeps the token sequence tractable while preserving fine-grained detail. It adapts across degradation types by predicting the reliability of internal layer features. Practitioners could use this as an alternative backbone to latent-diffusion restoration pipelines when detail fidelity matters more than leveraging existing T2I generative priors.

arXiv · cs.CVBuildable

Steering the Flow: Inverting Face Recognition Models via Gradient-Guided Flow Matching

Attackers can now reconstruct a stranger's face just by peering inside a face-recognition AI.

Face recognition systems convert your face into a compact numeric code; a 'model inversion attack' tries to reverse that process and recreate an actual face image from the code, which is a real privacy risk. Older attack methods produced unstable, unreliable results. This new method first trains a general-purpose model of 'what human faces look like' using a technique called flow matching, then in a second step nudges that generator step-by-step toward a specific target person by borrowing gradient signals directly from the face-recognition model itself, gradually ramping up how strongly it pushes toward the target. This matters because it reveals a security weakness in biometric systems, useful for researchers testing and hardening face-recognition privacy.

Technical view

SFMI is a two-stage white-box model inversion method: stage one pretrains an unconditional flow-matching prior over the face manifold as a robust generative prior; stage two, a Progressive Guidance Scheduler, injects time-dependent, identity-specific gradients obtained by backpropagating through the target recognition model during sampling, steering the generation trajectory toward the target identity. This reframes inversion as trajectory-steering rather than indirect/stochastic guidance, and could be used by security researchers to red-team face-recognition systems or as a baseline for measuring privacy leakage defenses.

arXiv · cs.CVConceptual

Revisiting Classifier-Free Guidance Methods in Latent Diffusion Models

None of the fancy tricks for improving AI image generation actually beat the classic method.

When AI image generators create a picture from a text prompt, a technique called classifier-free guidance (CFG) helps push the output to actually match what was asked for. Over the years, researchers have proposed many newer, fancier variants of CFG, but they were mostly tested on older AI models using metrics that only judge how nice the image looks — not whether it truly matches the prompt's meaning. This paper retests eight of these tricks on today's more advanced image generators, checking both quality and prompt-accuracy. The finding: none of the newer tricks reliably beat plain old CFG, and the one that seemed best (called APG) only wins by a margin so small it might just be noise. This matters as a caution against assuming a shinier new technique is automatically better.

Technical view

The study benchmarks eight training-free CFG-family inference techniques — many originally validated on older U-Net diffusion models — on two open-weight rectified-flow transformer diffusion models, under a fixed per-model protocol using three compositional-alignment (text-image correspondence) benchmarks rather than isolated image-quality metrics. No method consistently improves over vanilla CFG; APG achieves several nominal best scores but gains often fall within estimated evaluation uncertainty. Practitioners choosing an inference-time guidance method for modern rectified-flow models should re-validate any CFG variant against alignment benchmarks rather than assuming prior U-Net-era results transfer.

arXiv · cs.CVRunnable

Calibration-Free Vehicle Speed Estimation: A Monocular Keypoint-Template Approach

Clocking a car's speed from a single dash-cam video, no ruler or camera setup required.

Measuring how fast a car is moving from video usually requires calibrating the camera against known road markings or measurements beforehand. This system skips that step by fitting a detailed 36-point outline (like a wireframe) of a typical vehicle shape onto the car in each video frame, then using that geometry to work out, frame by frame, how pixels on screen correspond to real-world distances. A trained detector locates those 36 points automatically, and the car's speed is calculated either by tracking the points directly or by measuring fine-grained pixel motion across the image. Tested on hundreds of real traffic videos at speeds from 30 to 100 mph, it worked reliably without needing any special camera setup. This matters because it could let ordinary roadside or overhead cameras be used for traffic speed enforcement or monitoring cheaply.

Technical view

The method fits a 36-keypoint vehicle template per frame, using a YOLO-based keypoint detector trained on diverse datasets, and updates a homography matrix each frame to project 2D pixel displacement into metric space without requiring roadway calibration objects. Two estimation strategies are compared — keypoint-only tracking versus warped optical flow with dense spatial aggregation — validated on 400+ clips from the VS13 and BrnoCompSpeed datasets across 30-100 mph, with the warped optical flow variant achieving the lower MAE. Practitioners could deploy the keypoint detector and homography-update pipeline on existing uncalibrated roadside or overhead cameras for traffic speed monitoring.

arXiv · cs.AIBuildable

GRIP: Grounded Reasoning via Information-Restricted Premises

Starving an AI's memory of easy answers forces it to actually read its sources.

Retrieval-augmented AI systems fetch documents to help answer a question, but often the AI's language model is so powerful that it just answers from what it already 'knows' internally and barely looks at the fetched evidence — a problem the authors call query dominance. GRIP fixes this by choking the evidence channel: the AI still sees the full question clearly, but the retrieved documents get squeezed through a narrow, noisy bottleneck that only lets through information the question alone couldn't supply. That forces the model to genuinely lean on the evidence for anything it can't already guess. Across five reasoning tests this cut a measure of 'ignoring the evidence' by about 30-fold and reduced made-up answers (hallucinations) by 73%.

Technical view

GRIP introduces deliberate capacity asymmetry in a RAG architecture: the decoder retains full-dimensional access to the query while retrieved-evidence representations pass through a stochastic (variational-style) information bottleneck, compelling that channel to encode only residual, query-independent information. Evaluated on five reasoning benchmarks against iterative RAG baselines, it drives a query-latent mutual-information diagnostic from 14.8 to 0.47 bits (~30x reduction) and cuts hallucination rate by 73%. Residual-alignment analysis shows bottlenecked evidence embeddings occupy subspaces less correlated with the query representation, supporting the claim that the model is forced to draw on genuinely new information. Practitioners building RAG pipelines could adopt the bottleneck module as a drop-in regularizer to diagnose and reduce query-dominance failures.

arXiv · cs.CRConceptual

Topological Attribution Distance (TAD): Revealing Segment-Level RAG Influence on LLM Output Geometry for Incident Log Analysis

Mapping the 'shape' of AI answers to trace which log line actually caused a cybersecurity verdict.

When an AI helps a security team analyze incident logs, analysts need to know exactly which piece of evidence led to its conclusion — especially if the AI's autonomous decisions turn out wrong. The problem is that many incident log entries look nearly identical, so current tools struggle to tell which specific segment of text actually swayed the model. This paper proposes Topological Attribution Distance (TAD), a way of examining the geometric 'shape' of how the AI's internal representations shift when different log segments are fed in, to pinpoint influence more precisely than existing methods. The goal is building trust and traceability so security teams can audit AI-driven decisions after the fact.

Technical view

The work targets evidence attribution in retrieval-augmented LLM pipelines used for cybersecurity incident-log analysis, where existing attribution methods fail to discriminate between highly similar log segments. TAD proposes measuring segment-level influence via the topology/geometry of the LLM's output representation space rather than surface-level similarity, aiming to capture holistic structural shifts induced by each retrieved segment. The stated goal is a provenance-tracking technique robust enough for agentic AI systems where decision-chain traceability is operationally critical. The abstract is truncated before methodological or quantitative details are given, so specifics of the topological measure and validation results are not yet available.

arXiv · cs.LGConceptual

Beyond $L_2$: Generalizing Abductive Latent Explanations to Diverse Prototype-Based Architectures

Teaching 'explainable AI' proofs to work on curved, non-Euclidean brains, not just flat ones.

Some neural networks are built to be interpretable by design, using 'prototypes' — reference examples the network compares new inputs to, like matching a photo against a mental template. A method called Abductive Latent Explanations (ALE) can mathematically guarantee why the network reached its answer, but only when the network's internal space behaves like ordinary flat, ruler-measured (Euclidean) geometry. Many cutting-edge networks instead use curved or probabilistic spaces — spherical distances, bell-curve-shaped uncertainty regions, or compressed projections — which broke ALE entirely. This paper extends ALE's formal guarantees to work across those more exotic geometries, closing the gap so more modern interpretable architectures can get the same rigorous explanations.

Technical view

Abductive Latent Explanations (ALE) generate formal, provably-bounded explanations for prototype-based classifiers by computing tight bounds on latent-space distances, but prior formulations assumed Euclidean metrics. This work generalizes ALE to non-Euclidean prototype architectures — spherical metrics, Gaussian-density representations, and dimensional-projection spaces — by systematically deriving distance/bound formulations appropriate to each geometry. The contribution is a framework extension rather than a new architecture, enabling formal safety-and-readability guarantees on state-of-the-art prototype networks that previously fell outside ALE's applicability. Researchers working on interpretable-by-design models could apply these derivations directly to add certified explanations to non-Euclidean prototype models they already use.

arXiv · cs.CVRunnable

TRACE-Bench: Decomposing and Diagnosing Multi-Reference Image Generation

A benchmark that breaks 'combine these reference images' AI tasks into buildable-block operations to diagnose failures.

Newer AI image generators can take several reference images — say a person, an outfit, and a background — and combine them into one new picture, but testing how well they do this has been messy because benchmarks group tasks into vague labels like 'subject composition' that don't capture the real complexity. TRACE-Bench instead breaks any such task down into four basic building-block operations (anchoring, disentangling, applying, and composing), so any prompt can be described as a formula combining these operators. By controlling how many operators a task uses, researchers can precisely dial up complexity and see exactly where a model starts to fail. This gives a much more diagnostic, apples-to-apples way to compare and improve multi-reference image generation systems.

Technical view

TRACE-Bench reframes multi-reference image generation evaluation around four atomic operators — Anchor, Disentangle, Apply, and Compose — allowing any prompt to be expressed as a compositional formula whose 'slot count' quantifies structural complexity. The benchmark comprises ~1,600 evaluation cases spanning slot counts 1–8, derived from 631 formula templates, replacing task-type-based benchmarks that suffered fragmented coverage and uncontrolled difficulty. This design lets practitioners isolate exactly which operator or complexity level causes model degradation rather than getting only aggregate task scores. Anyone building or fine-tuning unified multimodal generation models can use it as a controlled diagnostic suite to target specific capability gaps.

arXiv · cs.AIBuildable

LAVA: Logic-Aware Validation and Augmentation Framework for Large-Scale Financial Document Auditing

An AI pipeline that reads messy financial documents and checks every number and rule like an auditor would.

Auditing financial documents — payroll, tax filings, loan applications — is hard to automate because the paperwork comes in wildly different layouts, uses dense business jargon, and must obey specific rules that a simple text-reading AI often mangles. LAVA is a four-step pipeline built on multimodal AI (models that read both text and layout/images) that first finds the relevant rule for a document, then extracts information while preserving its original layout, adds helpful extra context, and finally runs auditable math and logic checks to verify everything adds up. Because each step is traceable, an auditor can see exactly which rule was applied and why a document passed or failed, which matters enormously when the AI's mistakes could mean real financial or legal consequences.

Technical view

LAVA is a modular, backbone-agnostic pipeline for financial document auditing structured as four stages: document-rule retrieval, layout-preserving multimodal information extraction, auxiliary metadata enrichment, and symbolic/arithmetic verification for auditability. It targets production settings (payroll, tax compliance, loan underwriting) where heterogeneous layouts and embedded business rules break conventional extraction pipelines, emphasizing fine-grained error attribution and reproducible, traceable execution end-to-end. Being backbone-agnostic, it's designed to plug in different multimodal LLMs as the underlying extraction/reasoning engine. Teams building document-processing systems for regulated industries could adopt the four-stage decomposition directly to add rule-grounding and verifiable arithmetic checks to existing extraction pipelines.

arXiv · cs.LGConceptual

On the Principles Behind Neural Network Optimizers

Why Adam, the AI training algorithm everyone uses, actually works — and when it quietly breaks.

Almost every large language model today is trained using an optimization algorithm called Adam, which decides how to nudge millions of internal numbers toward better performance, yet surprisingly nobody fully understood why it works so well or when it might fail. This thesis shows there's a tipping point: with the right settings tuned to how much data you process at once, Adam reliably improves, but with certain other settings it can spiral out of control instead. It also explains why Adam beats the older, simpler method called SGD specifically on Transformer models (the architecture behind ChatGPT-style AI): the mathematical 'curvature' of the training landscape naturally organizes itself into distinct blocks, and Adam's design happens to handle that block structure especially well. Understanding this helps researchers design better, more reliable training algorithms instead of relying on trial and error.

Technical view

This thesis provides theoretical grounding for Adam by establishing a problem-dependent phase transition: with batch-size-dependent hyperparameter choices Adam provably converges, while small-β2 regimes can induce divergence. It further explains Adam's empirical superiority over SGD on Transformers via Hessian analysis, showing the loss curvature evolves toward a near-block-diagonal structure with strong inter-block heterogeneity as training progresses, and proves this structure is what makes Adam's per-coordinate diagonal preconditioner effective (unlike SGD's isotropic updates). The block-diagonal Hessian structure is traced to consecutive multiplications of large weight matrices characteristic of deep/attention architectures. This gives optimizer designers a concrete mechanistic target — block-heterogeneous curvature — to exploit when proposing Adam alternatives or scheduling hyperparameters like β2 relative to batch size.

arXiv · cs.CVBuildable

Binarized High-Efficiency RAW Video Restoration and Beyond

Shrinking video-restoration AI to near-binary math while losing almost none of its picture quality.

Cameras' raw sensor data is noisy and needs 'restoration' processing before it looks good, and doing this for video in real time on small devices is expensive. Binary neural networks — which represent numbers using essentially just 1s and 0s instead of full precision — can make image processing far cheaper, but they've struggled with video because they lose track of how frames relate to each other over time and mishandle the statistical spread of values. BinRVR solves this with a module that jointly models space and time efficiently in this ultra-compressed format, plus a smarter binarized building block that uses statistics from the original full-precision version to avoid quality collapse. The result cuts computation and parameters by about 96% while only losing about 4% of restoration quality — a huge efficiency win for deploying video enhancement on lightweight hardware.

Technical view

BinRVR is a binarized neural network for RAW video restoration that addresses BNNs' known weaknesses in temporal modeling and activation-distribution mismatch via two components: a Binarized Information Interaction Module (BIIM) that jointly captures spatial and temporal dependencies in the binarized domain, and a Distribution-Aware Binarized Convolution (DAB-Conv) that uses full-precision activation statistics to guide quantization and reduce error. Reported results show ~96% reduction in computation and parameters versus a full-precision baseline with only ~4% performance degradation. This positions BinRVR as a strong candidate for deploying low-level video restoration on resource-constrained edge devices, and the BIIM/DAB-Conv components could be reused in other binarized video-processing architectures facing similar temporal-modeling and quantization challenges.

arXiv · cs.CVBuildable

Beyond Uncertainty: Generalizable Failure Monitoring for Surgical Segmentation under Acquisition Degradation

Catching surgery AI's confidently wrong guesses before they mislead the surgeon.

AI systems that outline organs and tissue during surgery (segmentation) can fail silently — producing a wrong outline while still 'feeling' confident about it, especially when the camera image quality degrades. Most safety checks today only watch the AI's confidence level, which misses exactly these confident-but-wrong failures. TCSR-Monitor is a separate watchdog system that sits alongside the segmentation AI without needing to peek inside it, and instead checks whether the outlined shapes make anatomical sense, stay consistent from one video frame to the next, and match what you'd expect given the image quality. Tested on real surgical video with various simulated camera problems it hadn't seen before, it caught far more of these dangerous silent failures than confidence-based methods alone, which matters for keeping AI-assisted surgery safe.

Technical view

TCSR-Monitor is a post-hoc, model-agnostic failure monitor for surgical segmentation that combines confidence scores with observable shape plausibility, temporal-consistency, and image-quality cues, requiring no access to model internals or ground truth at deployment. It wraps a frozen segmentation model and includes a validation protocol specifically designed to test whether raised alarms remain credible under distribution shift. In leave-one-corruption-out evaluation on EndoVis 2017, it generalizes to unseen acquisition degradations and substantially outperforms confidence-only baselines at catching confident-but-wrong ('silent') failures. Its black-box, no-ground-truth design makes it directly deployable as a wrapper around existing production segmentation models for real-time surgical risk monitoring.

arXiv · cs.LGConceptual

Would this change your answer? Evaluating Explanations of LLM Behavior In The Wild with Counterfactual Experiments

Researchers test whether an AI's self-explanation actually predicts what it'll do next.

When an AI chatbot behaves oddly, researchers often ask it to explain itself — but how do we know if that explanation is actually true, rather than just a plausible-sounding story? This paper proposes a clever test: take the explanation, use it to predict how the model would behave if you tweaked the situation slightly (a 'counterfactual'), then check if the prediction holds. The team built an automated pipeline, CHIVE, that hunts for weird AI behaviors in real usage, edits the prompts that caused them, and collects thousands of these explanation-plus-evidence pairs. Surprisingly, they find that popular AI-interpretability tools don't necessarily help predict behavior better than just asking the AI directly — a humbling result for the field of 'looking inside' AI models.

Technical view

CHIVE (Counterfactual Hypothesis Investigation Via Edits) is an agentic pipeline that mines production LLM traffic for anomalous behaviors, generates counterfactual prompt edits, and evaluates explanations via counterfactual simulatability — whether an explanation lets you predict the model's behavior on related but altered inputs. The pipeline yields a large-scale naturalistic benchmark of explanation-counterfactual pairs, sidestepping the synthetic/toy-task limitation of prior interpretability evaluations. They use it to benchmark whether common interpretability techniques (e.g., feature attribution, chain-of-thought inspection) actually improve an agent's counterfactual prediction accuracy versus baseline self-report, finding surprising results that question current interpretability tooling's practical utility. Practitioners could adopt CHIVE as a standardized eval harness for any new interpretability method against real deployed-model behaviors.

arXiv · cs.CVBuildable

VicEdit: Learning to Edit Videos from Visual In-Context Examples

Instead of typing 'make it look like this,' you just show the AI a picture — it edits video that way.

Editing a video with text instructions is clumsy when you want a very specific texture, lighting style, or motion — words just can't capture it precisely. This paper flips the approach: instead of describing the edit in words, you show the AI an example — a reference image, a before/after image pair, or even a reference video clip — and it figures out what change you want and applies it. To teach a model this skill, the researchers built a massive dataset of 400,000 examples covering ten kinds of edits, automatically generated and quality-filtered. The resulting system, VicEdit, blends these visual hints with any text you still want to add, aiming to make video editing feel more like showing-and-pointing than writing precise prompts.

Technical view

VicEdit introduces 'visual in-context editing' as an alternative to text-only instruction-based video editing, accepting heterogeneous visual references — single images, image pairs, or video pairs — alongside text. The authors curate VicEdit-400K via an automated generation-and-filtering pipeline spanning ten edit task types, addressing the lack of large-scale supervision for this paradigm. The core technical contribution is a Modality-Adaptive Semantic (extraction) mechanism designed to pull editing intent out of visually heterogeneous reference formats and fuse it with textual context within a unified conditioning framework. This targets the well-known failure mode of text-only video editors on fine-grained texture/dynamics transfer, and the dataset itself is a reusable resource for future visual-conditioned video editing research.

arXiv · cs.LGBuildable

Le Critique: Privileged Value Functions for LLM Reinforcement Learning

A smarter 'coach' signal helps AI models learn from trial-and-error training faster and more precisely.

When AI models are trained using reinforcement learning (rewarding good outputs, penalizing bad ones), a key challenge is figuring out exactly which part of a long response deserves credit or blame — not just 'this whole answer was good,' but 'this specific step was good.' Popular methods like GRPO get around this by generating many attempts per question and comparing them, but that's slow and wasteful, especially when a few slow 'straggler' attempts hold up the whole batch. This paper revives an older idea — a learned 'critic' that scores each individual step — but makes it practical by feeding it privileged extra information during training that it won't have at test time, like a coach who's seen the answer key. The goal is training that's both more precise and less bottlenecked by slow outliers.

Technical view

The paper targets the sequence-level credit assignment and straggler-induced throughput/off-policyness problems in group-relative RL methods (e.g., GRPO) for LLM post-training. It proposes Privileged Value Functions (PVF), which inject additional task-relevant token-level signal into a learned critic during training (signal unavailable at inference), aiming to recover the token-level advantage estimation benefits of classic actor-critic RL without the usual infrastructure overhead that has made critics unpopular versus critic-free methods. This is a systems-plus-algorithms contribution relevant to anyone building RLHF/RLVR pipelines who wants finer-grained credit assignment and better rollout throughput than pure group-relative sampling provides.

SYS

Systems, OS & Low-Level

47 new
arXiv · cs.AIConceptual★ flagship

Quipu: A Governed Bitemporal Knowledge Graph Store

A database for AI-written knowledge that vets, timestamps, and audits every fact at the gate.

As AI agents start writing into knowledge graphs — big webs of interconnected facts — the databases holding them still assume a careful human curator, accepting whatever's written and cleaning up later. Quipu rebuilds the store around distrust of the writer: no fact gets in unless it passes a gate that checks what the database would look like after the write. Everything — the data, the trust labels, the approval verdicts, even the governance rules themselves — is 'bitemporal,' meaning it tracks both when something was true in the world and when the system learned it, so you can reconstruct any past state. Trust is organized by named sub-graphs whose combinations can only narrow authority, never secretly widen it, and the entire rulebook and audit trail live inside the same store as queryable facts. It matters because agent-generated knowledge needs provenance and governance baked in, not bolted on with dashboards.

Technical view

Quipu is an embeddable bitemporal knowledge-graph store that inverts four legacy defaults: (1) admission control via a gate whose predicates evaluate the pending post-state before any write commits; (2) full bitemporality over data, trust labels, verdicts, and the governing rules; (3) named graphs as the unit of authority/trust, composed under a lattice whose sole invariant is that composition never widens authority; and (4) reifying the governance spec Σ, the execution trace, and signed verdicts as first-class facts, so the compliance check T ⊨ Σ reduces to a query. Evaluation uses Census, a deterministic benchmark. Practitioners get an auditable, self-describing store where governance is queryable rather than middleware-enforced, suited to multi-writer agent workloads needing provenance and non-widening trust composition.

arXiv · cs.ARBuildable★ flagship

ESR-HGNN: Eliminating Semantic Redundancy for Efficient Mini-batch HGNN Inference

Speed up graph-AI predictions by not re-walking the same paths over and over.

Heterogeneous graph neural networks are AI models for data with many kinds of nodes and links — think a network with users, products, and reviews all connected. To run on huge graphs, they sample small batches by walking predefined 'metapaths' (patterns like user→product→user), but this walking causes lots of scattered, slow memory lookups and repeats many of the same traversals redundantly. This work spots that redundancy and eliminates it by storing traversal paths in a shared 'trie' — a tree structure that lets overlapping walks reuse a common prefix instead of recomputing it. That cuts the wasted repeated work and the random memory access that bottleneck inference. It matters because sampling, not the neural network math, is often the real speed limit when serving these models at scale.

Technical view

ESR-HGNN targets the metapath-based mini-batch sampling bottleneck in end-to-end HGNN inference, where irregular graph traversal induces heavy random memory access and semantic redundancy causes repeated traversals. The core idea is a redundancy-aware sampling paradigm using a metapath trie that reuses shared traversal prefixes across metapaths, eliminating duplicate walks and improving sampling efficiency. This directly reduces the dominant sampling cost, yielding faster mini-batch inference without changing model semantics. Practitioners serving HGNNs could integrate the trie-based sampler as a drop-in replacement for naive metapath sampling to cut inference latency and memory-access overhead on large heterogeneous graphs.

arXiv · cs.DBConceptual

Rerootable Hypertree Decompositions

Fixing a decades-old math trick for database queries so it can be rebuilt from any starting point.

Databases answer complex questions (like 'find all customers who bought X and reviewed Y') efficiently partly thanks to a mathematical structure called a hypertree decomposition — think of it as a smart tree-shaped map of how to break a big query into small, solvable pieces. A useful property would be 'rerootability': being able to pick up that tree map and re-flip it to start from any point, the way you can re-root a family tree from any person. This paper shows that current versions of these decompositions don't actually have that property reliably, traces the problem to deeper technical issues, and designs a new relaxed version that does support rerooting while staying computationally efficient. Experiments show this fix costs only a modest efficiency penalty — good news for eventually using these ideas in real database engines.

Technical view

The paper addresses rerootability — the ability to re-derive a valid decomposition rooted at any node from an existing one — in hypertree decompositions for conjunctive query answering, a property largely unexamined in the theory literature despite (or because of) practical relevance. The authors show standard 'normal form' decompositions, needed for tractability guarantees, actually obstruct rerootability, so they define a relaxed normal form that restores rerootability while preserving tractable evaluation. They connect this to 'projection-freeness' as a key structural property and provide experimental evidence quantifying the width increase (a proxy for query evaluation cost) incurred by the relaxed class versus standard decompositions. This is directly relevant to anyone building or theorizing about query optimizers that rely on hypertree-width-based planning.

arXiv · cs.DCBuildable

Bounded-State Restoration: Decoupling Local Restore Capacity from External LLM State

A trick to keep AI's 'memory' of long conversations light on local hardware without losing any of it.

When an AI model has a very long conversation or document loaded, it stores a huge amount of internal 'working memory' (called KV-cache) that's too big to fit in the GPU, so systems park extra memory elsewhere and pull it back when needed. The problem this paper identifies: even if you have plenty of storage for that overflow memory, actually reloading it and making it usable again requires its own separate chunk of fast local memory, and that requirement has been quietly ignored. They introduce Bounded-State Restoration, a technique that first checks how much of the stored memory can be reused without loading it all, then restores only a small moving 'window' of it at a time — like paging in a book a few pages at a chapter instead of loading the whole thing. This keeps the restoration process cheap regardless of how much total memory is being restored, which matters for making long-context AI conversations fast and affordable to serve.

Technical view

The paper isolates 'restoration working set' (RWS) as a distinct resource from cache retention capacity in hierarchical KV-cache systems (e.g., LMCache): the peak local staging memory needed during restoration, independent of how much total context state is retained upstream. They show empirically that in the current pinned upstream LMCache whole-plan restoration path, restoring GiB-scale KV states (1.956/7.823/15.646 GiB/rank) requires local L1 staging capacity that scales with total state size (2/8/16 GiB rungs), effectively coupling local memory cost to total retained context. Bounded-State Restoration (BSR) decouples discovery (probing the full reusable prefix) from materialization (installing state through a bounded window of at most W chunks), yielding O(W) peak local restoration memory independent of total reusable state size. This is directly actionable for anyone deploying long-context LLM serving with hierarchical/offloaded KV-cache — it targets a concrete memory bottleneck in systems like LMCache.

arXiv · cs.NIConceptual

Threat Aware Task Offloading and Caching for Secure UAV Assisted Vehicular Consumer Electronics

Drones help cars offload computing tasks safely, dodging both hackers and slow networks.

Modern cars are increasingly like rolling computers, running demanding apps (navigation, safety alerts, entertainment) that need fast processing, but a car's onboard computer isn't always powerful enough, so it 'offloads' work to nearby roadside stations or, in this paper, to drones (UAVs) hovering overhead. The catch is that wireless connections used for offloading can leak private data or get manipulated by attackers, especially as traffic and connections shift constantly. This paper designs a system that watches for suspicious communication patterns and factors that risk into its decisions about where to send each task and what data to pre-cache nearby, all while trying to keep response times low. It's essentially building a security-conscious traffic controller for the invisible computing work happening around and above your car.

Technical view

The paper proposes a UAV-assisted vehicular edge computing (VEC) architecture that jointly optimizes task offloading and spatiotemporal caching across roadside units (RSUs) and UAV edge nodes under threat-aware constraints. It formulates a security-aware uplink transmission model capturing information-leakage risk and anomalous communication patterns, feeding these risk estimates into adaptive offloading decisions, then poses a joint optimization minimizing end-to-end task latency while improving cache utilization under this threat model. This targets the intersection of edge-computing resource allocation and physical/communication-layer security in vehicular networks — relevant to researchers building secure VEC/UAV offloading schedulers or evaluating information-leakage-aware optimization formulations.

arXiv · cs.ARConceptual

ETHEREAL: A 25.6-$μ$s/inf. Low-latency Event-driven Graph-neural-network Processor for High-resolution Vision at the Edge

A new chip lets cameras that only 'see' motion process it in millionths of a second, at the edge.

Special cameras called dynamic vision sensors don't record full video frames like normal cameras — they only fire off tiny 'events' the instant something changes in the scene, which is extremely fast and efficient, similar to how our eyes' retinas work. Making sense of this stream of scattered events requires a particular kind of AI (event-driven graph neural networks) that's good at handling sparse, irregular data, but until now, no hardware chip was specifically built to run this efficiently. ETHEREAL is the first chip purpose-built for this job, combining a specialized processing engine with a clever memory layout designed for these scattered, time-and-space-stamped events. The payoff is vision processing fast enough (25.6 millionths of a second per inference) for things like drones or robots needing near-instant reactions, using far less power than running the equivalent AI on a general-purpose chip.

Technical view

ETHEREAL is presented as the first dedicated hardware processor for event-driven graph neural networks (EV-GNNs), which process sparse spatiotemporal event streams from dynamic vision sensors (DVS) for low-latency edge vision. The architecture combines a neighbor-parallel spline-convolution engine (handling the dense-regular compute of graph convolutions) with a split-2D/3D memory hierarchy incorporating a novel spatiotemporal event-caching mechanism to handle the sparse-irregular memory access patterns inherent to graph-structured event data. The headline result is 25.6 microseconds per inference, targeting sub-millisecond edge-vision latency budgets that neither conventional frame-based vision pipelines nor general-purpose GNN accelerators comfortably meet. This is a hardware/silicon contribution — of direct interest to chip designers or roboticists building ultra-low-latency perception pipelines around event cameras, though it requires custom silicon rather than being reproducible in software alone.

arXiv · cs.NIConceptual

Array-Based Molecular Pulse Encoding for Neuro-Spike Communication in Intra-Body Nano-networks

Tiny medical robots could bridge severed nerves by 'talking' in coded chemical pulse patterns.

When neurons in the body are damaged, one futuristic fix is to use engineered nano-machines that sit at the break and relay the signal across, mimicking how real neurons communicate using chemical 'spikes.' Real neurons encode information partly by exact timing — how fast the spikes fire and the gaps between them — but that requires the sender and receiver machines to have their clocks precisely synchronized, which is very hard for tiny, resource-poor devices to pull off. This paper proposes a workaround: instead of relying on exact timing, encode information in the pattern or arrangement of different types of molecular pulses (like a sequence of different-flavored signals) that a receiver can decode just from noticing the sequence, without needing tight time-syncing. It's an early building block for engineering nano-scale communication networks that could one day help repair nerve damage.

Technical view

The paper addresses a synchronization bottleneck in neuro-spike molecular communication systems designed to bridge severed neural connections using nano-machine relays. Conventional temporal spike-rate/inter-spike-interval encoding demands tight time synchronization between transmitter and receiver nano-machines, infeasible given their resource constraints. The proposed scheme instead encodes information via the arrangement (ordering/combination) of distinct molecular pulse types emitted in an array, requiring only symbol-level (not exact temporal) synchronization, with symbols distinguished by pulse-type composition rather than precise timing. This is a physical-layer communication-theory contribution to the intra-body/molecular communication field, relevant to researchers modeling channel capacity, error rates, or detector design for reduced-synchronization nano-network schemes.

arXiv · cs.LGRunnable

rl-triton: High-Performance Triton GPU Kernels for Reinforcement Learning Credit Assignment

One math trick makes seven different reinforcement-learning credit-scoring algorithms run 5x faster on GPUs.

When an AI agent plays a game or controls a robot, it needs to figure out which past actions actually led to a good or bad outcome — this is called 'credit assignment,' and there are seven popular mathematical recipes for computing it (with names like GAE and TD-lambda). The researchers noticed that all seven are secretly the same underlying calculation in disguise, just a chain of numbers each depending on the one before it. They rewrote that shared calculation as a 'parallel scan,' a technique that lets a GPU compute a long chain of dependent steps almost all at once instead of one-by-one, and hand-tuned the low-level GPU code (using a language called Triton) to do it efficiently. The payoff is that training loops for reinforcement learning — which are often bottlenecked on exactly this kind of computation — get up to 5.7x faster with no change in the results.

Technical view

rl-triton unifies seven RL return/advantage estimators (GAE, V-Trace, Retrace(λ), TD(λ), discounted returns, eligibility traces, episodic prefix sums) under a single associative-scan formulation of a first-order linear recurrence, solvable in O(log T) parallel depth instead of O(T) sequential steps. Each algorithm supplies its own recurrence coefficients via algorithm-specific fused Triton kernels sharing one associative operator, with explicit, algebraically-verified handling of episode termination vs. truncation boundaries. Benchmarked against a vectorized torch.compile baseline, it achieves 1.6–5.7x full-call speedup on massively parallel RL workloads. Practitioners can drop it in as a kernel-level replacement wherever these return estimators currently form a training bottleneck, and the associative-operator framework offers a template for fusing other RL recurrences.

arXiv · eess.SYBuildable

Communication Reduction via Semantic-Based Encoding in DMPC Using LSTMs

Robots learn to compress their chatter into tiny messages so swarms don't jam the network.

Groups of mobile robots that plan their movements together (distributed model predictive control) normally need to constantly radio full details of their plans to each other, which can flood the wireless network. This work trains a neural network — specifically one built from LSTMs, a type of network good at handling sequences of data over time — to squash each robot's message down into a smaller, encoded version before sending, and to reconstruct a close approximation of the original message on the receiving end. Think of it like each robot speaking in a compact shorthand that its teammates have learned to decode. Tests with real robot formations show this compressed communication holds up well even when the network is overloaded, and the same trained network can be reused across different planning horizons without retraining, which saves engineering effort.

Technical view

The paper applies LSTM-based encoder-decoder networks within a distributed MPC optimization loop, where each agent transmits a learned low-dimensional latent encoding of its planning message instead of the raw data, and receivers decode it back to a usable reconstruction. Evaluated on mobile robot formation-control tasks under constrained bandwidth, the compressed scheme maintains control performance and remains reliable in bandwidth regimes that saturate full-message communication. Notably, the LSTM-based approach generalizes across different prediction-horizon lengths without retraining, suggesting the learned encoding captures horizon-invariant structure in the DMPC messages. This offers a practical building block for bandwidth-constrained multi-robot systems using semantic/learned compression rather than classical quantization or sparsification.

arXiv · cs.DCConceptual

Counting in Population Protocols on Graphs

Anonymous robot swarms can now count themselves fast, just by randomly bumping into each other.

Imagine a huge swarm of identical, nameless devices (like tiny sensors) that can only interact in pairs when randomly picked, and none of them know how many total devices exist — this is the 'population protocol' model. This paper solves the problem of getting the whole swarm to collectively figure out its own size, but now allowing the devices to be connected in a specific network graph rather than assuming everyone can reach everyone. Their trick is to use random coin-flips generated during each random pairwise meeting, plus a clever way to estimate the swarm's approximate size on a logarithmic scale, all while relying only on how fast information can spread across that graph. The result is a protocol that counts accurately using very little memory per device and converges provably fast, which matters for designing self-organizing systems like sensor networks or synthetic biological circuits where no central coordinator exists.

Technical view

The paper studies size-counting in population protocols where agents interact over an arbitrary graph G via a random-edge scheduler, with one interacting party designated initiator per step to break symmetry among anonymous, identifier-free agents. Their protocol uses Õ(n) states per agent and stabilizes with high probability in O(B(G)·log²n + L(G)·log n) interactions, expressed in terms of the graph's broadcast time B(G) and load-balancing time L(G), generalizing prior complete-graph results to arbitrary topologies. Key technical contributions are novel sub-protocols for sampling independent random bits from interactions and for approximating log n up to additive error using only local pairwise exchanges. This gives a template for adapting other population-protocol primitives (e.g., majority, leader election) to graph-structured rather than fully-mixed populations.

arXiv · cs.DCConceptual

Optimal Adaptive Multi-Valued Byzantine Agreement

A faster way for many computers to agree on a big shared value, even if some are lying.

Byzantine Agreement is the classic problem of getting a group of computers to settle on one shared answer even though some of them might be malicious or faulty and actively try to sabotage the process — it's foundational to blockchains and fault-tolerant systems. Older protocols assumed the worst-case number of bad actors and paid a heavy communication cost regardless of how many were actually misbehaving; a recent line of work showed you can make the cost scale with the actual number of cheaters rather than the worst case, but only for agreeing on a single yes/no bit. This paper extends that idea to let the group agree on a much richer value — a whole multi-bit message, like a full block of data — while keeping the same efficient, adaptive cost. That matters because real systems (blockchains, distributed databases) need to agree on rich values, not just single bits, and doing so efficiently reduces the overhead of running these networks at scale.

Technical view

Building on Constantinescu et al.'s adaptive binary Byzantine Agreement protocol achieving Õ(n + t·f) message complexity and Õ(f) rounds (where f ≤ t is the actual, not worst-case, number of corrupt parties), this work extends the framework to L-bit multi-valued agreement while preserving adaptive complexity and optimal resilience (t < n/2 synchronous). The construction combines the binary-agreement scaffolding with novel value-dissemination and validation techniques parameterized by a security parameter κ, avoiding the naive approach of running L independent binary instances. This yields the first adaptive, near-optimal-complexity multi-valued BA protocol, directly relevant to practitioners building high-throughput consensus layers (e.g., blockchain validators) who need to agree on full blocks rather than single bits.

arXiv · cs.CRBuildable

CryptDough: A Unified Analytics Engine for Secure Multiparty Computation

A single toolkit lets rival companies crunch each other's secret data together without peeking.

Secure multiparty computation (MPC) lets several organizations that don't trust each other jointly compute something — like a combined statistic — from their private data, without any of them revealing their raw data to the others. Until now, most MPC tools only handled one kind of security assumption or one kind of workload (say, just database-style queries, or just machine learning), forcing companies to stitch together different specialized systems. CryptDough is a single engine that handles relational data, time-series data, and ML inference all together, under multiple different trust/security models, using a layered software design so each layer can be swapped or extended independently. It also introduces 'virtual vectors,' a programming abstraction that lets developers write simple, ordinary-looking single-threaded code while the system automatically handles the messy parts — networking, parallelism, memory — behind the scenes. This matters because it could make privacy-preserving data collaboration (e.g., between hospitals or banks) far more practical to build and deploy.

Technical view

CryptDough is a general-purpose MPC analytics engine supporting cross-domain workloads (relational analytics, time-series processing, ML inference) under multiple threat models within one system runtime, contrasting with prior single-threat-model or single-workload MPC systems. Its architecture uses hierarchical, progressively-lowered abstraction layers for modularity, and introduces 'virtual vectors' — a programming abstraction letting developers author single-threaded-looking code while the runtime handles communication scheduling, parallelization, and memory management transparently across all layers. This design targets extensibility (new protocols/threat models can be added at appropriate abstraction levels) and could serve as a reference architecture for teams building unified privacy-preserving data pipelines rather than bespoke per-workload MPC stacks.

arXiv · cs.NIBuildable

Completion-Path Credits: Multi-Resource Control for Scale-Up Fabrics

Smarter traffic tickets for GPU networks stop tiny operations from getting stuck behind big ones.

When many GPUs are wired together in a 'scale-up fabric' to train huge AI models, they send both bulk data transfers and tiny operations (like atomic updates or quick notifications) over the same shared connections. Traditional traffic-control systems only track how many bytes are flowing, which works fine for big transfers but badly represents small operations whose cost isn't really about byte count — it's about hogging a specific limited resource, like the chip that processes atomic operations. SemaCredit fixes this by tracking a whole vector of separate resource 'budgets' — one for memory, one for atomics, one for response handling — and only lets an operation proceed if it has enough of each specific resource it actually needs, releasing each piece back exactly when that stage finishes. In simulated tests, this dramatically cuts the worst-case delays for small operations (over 50% improvement) during heavy atomic traffic, without hurting normal big-transfer performance, which matters for keeping large AI training clusters running efficiently.

Technical view

SemaCredit is a receiver-side admission controller for scale-up GPU/accelerator fabrics that replaces single byte-denominated credit pools with a vector of target-resource credits (HBM, Atomic engine, response-injection stage), admitting each remote-memory operation only against the specific resources it demands and returning each credit component independently as its corresponding pipeline stage completes. Evaluated in a deterministic event simulator with multipath queues, 8 HBM partitions, a serialized Atomic engine, and a response engine, it matches a strong per-resource byte-credit baseline on HBM-hotspot traffic while cutting small-operation P99 latency by 52.4% under Atomic contention and 10.2% under response incast. This is a concrete fabric-controller design practitioners building UALink/scale-up interconnect stacks could implement to better isolate small atomic/notification traffic from bulk tensor transfers.

arXiv · cs.NIBuildable

Predict Before Replay: Joint FEC and Flight Control for Reliable Scale-Up Links

AI chip networks predict incoming error bursts and brace for them before data even arrives.

Ultra-fast links between AI chips send tiny data packets ('flits') at blistering speeds, and when errors happen they usually come in bursts (many corrupted symbols in a row) rather than isolated single errors. The normal fix-up process — error-correcting codes plus 'replay' (resending lost data) — reacts too slowly at these speeds: by the time a corrupted packet is noticed, later packets have already been sent and now also need to be resent, turning one glitch into a much bigger pile of retransmissions. PREFACE tackles this by watching the pattern of small, already-corrected errors and using a statistical filter (a two-state Bayesian model) to predict whether a burst of worse errors is coming, then proactively strengthens the error-correction and limits how many packets can be 'in flight' unacknowledged at once. In simulation, this meaningfully boosts data throughput, cuts worst-case latency in half, and reduces wasted retransmissions, which adds up to notably faster collective communication (like the AllReduce operations used to sync huge AI training runs across many chips.

Technical view

PREFACE is a pre-FEC (forward error correction) controller targeting temporally-correlated burst errors on scale-up accelerator links, using a two-state Bayesian filter over corrected-symbol observations to infer a posterior probability of an imminent burst on the next flit, then jointly tunes FEC strength and an outstanding-flit cap in response. Implemented in ns-3 against UALink 200G 1.0's publicly specified replay semantics, it improves goodput by 10.52%, cuts P99 latency by 50.75%, reduces replay volume by 47.52%, and speeds up modeled ring AllReduce by 13.1–27% (range cut off in abstract). This gives link-layer designers a concrete, simulator-validated joint FEC/flow-control policy that could be implemented in real UALink-class controllers to reduce burst-error amplification during large-scale distributed training.

arXiv · cs.NIBuildable

LoRIS: LoRaWAN-based IoT Platform for Sustainability Monitoring in Hotels

Battery-powered sensors quietly track hotels' water, power, and waste to make sustainability claims real.

Hotels are surprisingly big contributors to carbon emissions, water use, and waste, but they mostly track their environmental impact through slow, manual, imprecise record-keeping rather than real measurement. LoRIS is a sensor network built on LoRaWAN, a long-range, low-power wireless technology, that scatters small battery-run sensors throughout hotel properties to automatically measure things like energy and water use, room conditions, and how guests actually behave, feeding it all into a dashboard. The system is specifically designed around real hotel constraints: it has to work with hotels' locked-down IT setups, respect guests' privacy, be easy to install and remove without tearing up walls, and run for years on one battery. It also builds in privacy protections from the start (choosing what to sense carefully and encrypting data end-to-end), which matters because it lets hotel chains finally back up their sustainability reports with real, high-resolution operational data instead of estimates.

Technical view

LoRIS is a LoRaWAN-based IoT sensing platform following the standard LoRaWAN reference architecture (end devices → gateways → network/application server), purpose-built for hospitality deployment constraints: restrictive hotel IT/network policies, multi-year battery-only operation, rapid non-invasive installation/removal, and guest privacy. It measures resource consumption, environmental conditions, and guest behavior across geographically distributed properties, with privacy-by-design principles governing sensing-modality selection and physical deployment zoning, plus end-to-end encryption from sensor to dashboard. This targets a concrete gap — hospitality sustainability reporting currently relying on coarse manual data — and offers a deployable reference architecture practitioners could adapt for other constrained commercial IoT/ESG-monitoring settings.

arXiv · cs.DCBuildable

Generalizing and accelerating consistency checking for non-transactional distributed storage systems

A faster judge for whether your distributed database kept its promises.

When companies build distributed databases spread across many machines, they need to test that operations look consistent to users even though data is copied everywhere. This paper improves a classic algorithm (Wing-Gong) used to check 'linearizability' — basically, does the system behave as if all its operations happened one at a time in a sensible order. The trick is generalizing the checker so it can also verify other, looser consistency rules that real systems like ZooKeeper actually promise, plus custom rules specific to each system. Testing on 8 real distributed storage systems shows the new checker catches more real bugs and runs up to 370x faster.

Technical view

The authors generalize the Wing-Gong (WG) linearizability-checking algorithm to support arbitrary non-transactional consistency models beyond strict linearizability, including ZooKeeper's ordered sequential consistency and system-specific ordering constraints derived from a system's spec. This lets testers encode additional partial-order constraints on operation histories rather than relying solely on real-time ordering, reducing false negatives when a system only claims weaker guarantees. Evaluated on 8 distributed storage systems, the generalized checker is easier to configure for system-specific guarantees, aids debugging of consistency violations, and achieves up to 370x speedups over standard WG-based checking, with better scalability to larger histories.

arXiv · cs.DSConceptual

A Black-Box Workload Barrier for Exact Girth via Multi-Scale Nearest-Source Estimation in CONGEST

Proving there's a hard mathematical wall to how fast networks can measure their own shortest loops.

In distributed computing, a network of computers (nodes) sometimes needs to figure out collective properties of its own connection graph, like the length of its shortest cycle (the 'girth') — imagine detecting the smallest loop in a giant web of routers, where each node only knows its immediate neighbors and can only send limited messages each round. Recent methods gave good approximate answers quickly, and this paper asks whether you can get the exact answer using a specific 'black-box' strategy of repeatedly sampling likely cycle members and adjusting based on what came back. Through a clever counting argument, they prove a hard limit: for a tricky family of networks, any such strategy can only succeed with a probability capped by how much total 'work' it performs, exposing an inherent tradeoff. It matters because it tells algorithm designers there's a real ceiling on how cheaply exact answers can be obtained this way, guiding where future research effort should go.

Technical view

The paper analyzes exact girth computation in the CONGEST model by isolating a specific black-box strategy — sequential adaptive calls to a nearest-source estimator on exchangeable source sets — and proves an implementation-independent workload lower bound. On a bounded-degree, log-diameter graph family with a unique girth-Θ(log n) cycle, they show via a permutation-rank argument that success probability is bounded by (3g_t/n_t)·E[Σ min{Q_j,k_j}], tying exactness to the total 'resolving work' (source counts × capacities) summed across adaptive calls. This formalizes why sublinear approximate girth algorithms can't be trivially converted to exact ones without paying substantially more communication, giving a concrete quantitative barrier for anyone designing or lower-bounding CONGEST distance/cycle-detection primitives.

arXiv · cs.DCConceptual

Brief Announcement: Fair Binding for Hidden-State Authorization in Byzantine SMR

Stopping sneaky double-spends when blockchain-style systems can't see what a command was actually allowed to do.

In systems where many possibly-malicious computers must agree on an ordered list of actions (state machine replication, used in blockchains and fault-tolerant databases), things get tricky when an action's validity depends on a hidden policy or resource the checking computers can't fully see — like an AI agent authorized to spend a budget the validators can't directly inspect. The authors point out that just proving 'this was allowed back when I checked' isn't enough, because two different requests could each show old proof of authorization for the same limited resource, and one needs to invalidate the other once used. Their fix requires two things: the order requests arrive must actually constrain the order they're committed in (already tackled by 'fair-ordering' protocols), and, crucially, whichever request commits first must automatically make later conflicting requests invalid rather than just recording history. This closes a subtle security gap that could otherwise let attackers double-spend a hidden, limited resource in agent-driven distributed systems.

Technical view

The paper addresses safe allocation of a hidden consumable resource in validated Byzantine SMR when validity depends on an evolving, unreconstructable policy/commitment state (e.g., agent authorization budgets) rather than the visible log. It decomposes correctness into two independent requirements: (1) arrival-order-preserving commit order, addressed by existing fair-ordering protocols, and (2) 'binding' — a committed first request must invalidate conflicting later ones, not merely leave a historical attestation of prior authorization. The contribution is showing requirement (2) is non-vacuous given current policy-commitment designs and proposing a binding mechanism so leader-controlled ordering can't be exploited to double-allocate hidden resources; relevant to anyone building agentic/LLM-authorized transactions atop BFT consensus.

arXiv · cs.ARBuildable

The Road Less Traveled: Congestion-Aware NoC Placement and Packet Routing for FPGAs

Teaching chip design software to route around network traffic jams inside next-gen FPGAs.

Modern FPGAs (reprogrammable chips) now build in dedicated data highways called Networks-on-Chip (NoCs) to move large amounts of data quickly across the chip instead of relying on scarce wiring. But this creates a new headache for the software that decides where to place a design's components and how to wire them (place-and-route): it now also has to avoid overloading those data highways (congestion) while still efficiently using regular wiring. This paper builds new placement and routing techniques, added to an open-source FPGA CAD tool, that specifically account for NoC traffic congestion when deciding where things go and how signals travel, aiming to avoid bottlenecks without wrecking other performance metrics like speed.

Technical view

The work extends an open-source FPGA CAD flow (VTR/VPR) to be NoC-congestion-aware: it adds a NoC link congestion cost term into the placement engine's objective function alongside existing wirelength/timing costs, plus new routing strategies to balance NoC bandwidth/latency utilization against conventional programmable-routing-resource usage of NoC-attached design modules. The approach targets link oversubscription specifically, aiming to reduce NoC congestion with minimal degradation to other standard PnR quality metrics. This is directly applicable to anyone using VTR-based flows to target commercial NoC-enabled FPGA architectures and wanting congestion-aware placement/routing out of the box.

arXiv · cs.NIBuildable

An O-RAN-Assisted MARL Approach for Dynamic Sidelink and Infrastructure Selection in V2X Communications

Letting AI agents in cell towers juggle car-to-car radio traffic so nobody's signal gets drowned out.

Future connected cars will talk directly to each other and to pedestrians' devices over short-range 'sidelink' radio channels, which is great for safety but risks interference — too many vehicles broadcasting can crowd out signals from vulnerable road users like cyclists. This paper uses Open RAN, a flexible cellular network architecture that gives operators a real-time, programmable view of the whole network, to run multiple AI agents (multi-agent reinforcement learning) that dynamically decide, for each device, whether to use direct vehicle-to-vehicle links or route through network infrastructure, and how to share resources. The AI agents learn from network feedback to balance the competing demands across the whole system rather than optimizing each connection in isolation, aiming to keep both vehicle and pedestrian communications reliable as traffic gets denser.

Technical view

The paper proposes a MARL-based resource-and-mode management system built on the O-RAN architecture for 6G V2X networks, jointly deciding sidelink vs. infrastructure-mediated communication mode and resource allocation across vehicles and VRUs. Multiple RL agents operate within O-RAN's control loops, leveraging its global network view and open interfaces to coordinate mode/resource decisions network-wide rather than per-link, explicitly targeting interference and starvation issues that existing pair-selection/resource-allocation-only approaches miss. This gives O-RAN researchers a concrete MARL xApp-style design pattern for holistic V2X mode selection that could be extended or benchmarked against centralized/heuristic baselines in sidelink simulators.

arXiv · cs.CLRunnable

Polaris: Learning to Generate Table Descriptions from Retrieval Feedback

Training an AI to write table summaries that actually help search engines find the right table.

When software needs to find the right database table to answer a question (say, for turning natural language into SQL), it often searches using auto-generated text descriptions of each table. The problem is that these descriptions are usually written by AI to sound fluent and readable, not to actually work as good search keywords, so retrieval quality suffers. Polaris flips this around: it generates several candidate descriptions per table, tests how well each one performs at actually retrieving the right table using a standard search method (BM25), and then trains the description-writing AI to prefer the versions that retrieve best, using a technique called Direct Preference Optimization. It also expands cryptic abbreviated column names first, so descriptions aren't tripped up by labels like 'cust_addr' — the result is descriptions optimized for being found, not just for reading nicely.

Technical view

Polaris fine-tunes an LLM to generate table descriptions optimized directly for retrieval effectiveness rather than fluency, using existing table-retrieval benchmarks' query-table relevance judgments as supervision: multiple candidate descriptions per table are generated, ranked by BM25 retrieval performance against benchmark queries, and turned into preference pairs for DPO fine-tuning. A preprocessing step expands abbreviated table/column names before generation to reduce vocabulary mismatch between queries and schema terms. This gives practitioners building NL2SQL or table-QA retrieval pipelines a recipe for repurposing off-the-shelf retrieval benchmarks as reward signal, without new human annotation, to train description generators that measurably improve downstream keyword-based table retrieval.

arXiv · cs.ARConceptual

RENESIS: Energy-Aware Synthesis of Adiabatic Logic from Irreversible Netlists

A compiler that turns ordinary chip logic into ultra-low-energy reversible circuits.

Normal computer chips waste energy every time they flip a bit and throw away the old value — that's basic thermodynamics at work. 'Adiabatic' or energy-recovery circuits try to get around this by never carelessly erasing information, recovering energy instead of dumping it as heat. RENESIS is a tool that automatically takes an ordinary chip design (the kind engineers already write) and converts it into one of these energy-recovering circuits, optimizing specifically for energy use rather than the usual targets of chip size or speed. It does this by treating the circuit as a kind of mathematical space and tracking, in detail, how information gets created, destroyed, or hidden as signals flow through — importantly, it treats 'reversibility' as a strict design rule to enforce, not just an abstract physics ideal, so the output is a real, checkable, low-power circuit blueprint.

Technical view

Renesis is an automated synthesis flow that transforms an ordinary irreversible netlist into a verified, technology-mapped adiabatic (energy-recovery) circuit, using energy consumption — rather than area/delay — as the primary optimization objective, and targeting one of eight known energy-recovery logic families. It represents the netlist in a vector-space formulation where simulation and justification are forward/reverse traversals with per-component linear cost, populating information-theoretic 'ledgers' that track switching, erasure, and observability at their natural Rényi orders to guide where logical reversibility must be preserved. Notably it treats reversibility as an enforced circuit-level constraint (not merely a thermodynamic lower bound like k_BT ln2), giving circuit designers a concrete, verifiable path from standard netlist inputs to energy-optimized adiabatic implementations.

arXiv · cs.AIBuildable

KernelArc: A Multi-Agent Framework for GPU Kernel Optimization

AI agents team up to write faster GPU code than humans can, and now top the leaderboard.

KernelArc is a system where multiple AI 'agents' each specialize in a different optimization strategy and work in parallel to rewrite the low-level code that runs on GPU chips, making AI models run faster. They coordinate by sharing only their conclusions (not their messy work-in-progress) and use a strict testing system to make sure any change actually helps before it's kept. Tested on NVIDIA's newest H100 and B200 chips, the resulting kernels — the tiny, hyper-optimized programs that do the actual math — beat other submissions on a public benchmark leaderboard. This matters because hand-tuning GPU code is a rare, expensive skill, and automating it could make AI systems cheaper and faster to run everywhere.

Technical view

KernelArc coordinates strategy-specialized agents via conclusions-only shared memory, a deterministic benchmark guard for reproducible measurement, and read-only cross-agent state with plateau-triggered drafting to avoid premature convergence. Evaluated on H100/B200 using SOL-ExecBench, outputs include custom BF16 GEMM kernels, static cuBLASLt Expert-API configs, fused MoE backward passes, shape-gated decoder fusion, native NVFP4 grouped-query attention, and paged prefill attention — ranking first on L1, L2, Quantization, and FlashInfer leaderboard categories as of July 30, 2026. Practitioners could adopt the conclusions-only coordination pattern for other multi-agent code-optimization tasks where naive shared-state causes agents to converge on the same local optimum.

arXiv · eess.SPBuildable

ECO-ID: Event-Camera based Optical System for Secure Multi-User Ultra-Low Latency Identification

Blinking LEDs invisible to normal cameras can ID dozens of people in microseconds.

ECO-ID is a way to identify multiple people or devices almost instantly using light instead of passwords, QR codes, or RFID chips. It uses a special 'event camera' that, unlike a normal camera, doesn't capture full frames — it only notices the exact microsecond a light gets brighter or dimmer, which is much faster and harder to snoop on remotely. Each person's location (which LED they're near) and precise timing pattern together spell out their unique ID, so no two users need to sync up with each other. Because it relies on visible light rather than radio signals, it's also harder for an attacker to intercept from outside the room, making it a promising building block for secure, instant sign-ins in things like shared AR headsets or smart venues.

Technical view

ECO-ID uses microsecond-resolution asynchronous event-camera sensing over visible light communication (VLC), combining spatial separation (disjoint LED subsets per user) with user-specific timing-delay codes to achieve collision-free multi-user identification without inter-user synchronization. Event-driven sensing avoids full-frame capture, reducing bandwidth/latency versus conventional cameras and shrinking the RF attack surface relative to RFID/NFC. The system supports rapid token verification with freshness and replay protection, suggesting applicability to secure device pairing or access control in AR/VR and IoT contexts where sub-millisecond, contactless authentication is needed.

arXiv · cs.CCConceptual

Superlogarithmic Gap Result for LCLs on Trees in Quantum-LOCAL

Quantum computers get no real speed edge over classical ones for a whole class of tree puzzles.

This is a theory paper about distributed computing — imagine many computers connected in a tree-shaped network, each only able to talk to its neighbors, trying to jointly solve a puzzle where the answer has to look locally consistent everywhere (a 'locally checkable labeling'). The question is whether letting each computer use quantum randomness helps them solve such puzzles dramatically faster than plain old classical computers. The authors prove that if a quantum-assisted approach can solve the puzzle reasonably fast, then a completely ordinary, quantum-free method can solve it almost as fast too (in logarithmic time). It's a 'no free lunch' result showing quantum tricks don't give a shortcut for this broad category of problems, which helps researchers know where quantum speedups can and can't be expected.

Technical view

The paper shows that on trees, any LCL solvable by an n^{o(1)}-round quantum-LOCAL algorithm (with bounded dependent distributions) can be simulated by an O(log n)-round deterministic classical LOCAL algorithm, via a rake-and-compress tree decomposition and local simulation of the dependent distribution on decomposition components. The corollary establishes a strict dichotomy: every tree LCL is either solvable in O(log n) deterministic rounds or requires n^{Ω(1)} rounds even with quantum-LOCAL, closing the gap between constant/log-round and polynomial-round complexity classes for this setting. This gives distributed-computing researchers a general derandomization/no-speedup tool for ruling out quantum advantage on tree-structured LCL problems.

arXiv · cs.NIConceptual

Expanding Access, Exposing Risk: A Short Study of Exposed Starlink Hosts

Starlink internet users are running riskier, more out-of-date software than everyone else online.

Researchers scanned the internet to see how secure the computers and devices connected via Starlink (Elon Musk's satellite internet) are compared to regular internet users. They found Starlink-connected hosts are more likely to be running old, vulnerable software and outdated network protocols than non-Starlink hosts — essentially, more unlocked doors. The problem is worse in places like Latin America, Southeast Asia, and Eastern Europe, where Starlink is often filling gaps left by weak traditional infrastructure. This matters because as satellite internet expands access to remote or underserved regions, it may be quietly importing new security risks alongside the connectivity, which policymakers and security researchers need to address.

Technical view

Using internet-wide measurement scanning, the authors compare the software/protocol freshness of Starlink-announced IP space against non-Starlink hosts, finding elevated rates of outdated OS versions and vulnerable network protocol implementations on Starlink hosts, with pronounced regional skew toward Latin America, Southeast Asia, and Eastern Europe. The short paper is primarily a measurement finding rather than a proposed fix, positioning it as a call to action for internet measurement and policy communities to investigate causes (e.g., CGNAT/terminal configuration defaults, underserved-region deployment patterns) and mitigations. Replicable via standard scanning tools (e.g., Censys/Shodan-style banner grabs) cross-referenced against Starlink's published ASN/IP ranges.

arXiv · cs.DCConceptual

HAPS through the Lens of Satellites and UAVs: A Function-Level Perspective on the Emerging High Altitude Economy

The 'third tier' between drones and satellites has flown for years — but has it actually done much yet?

High-Altitude Platform Stations (HAPS) are aircraft or balloons that hover in the stratosphere, 17-27 km up — higher than drones, lower than satellites — and people have long argued they could combine the best of both for things like earth-watching cameras, GPS-like navigation, and internet relay. This paper is a reality check: instead of trusting the hype, the authors go through nineteen proposed uses one by one and ask, strictly, has this actually been proven with real flight data at altitude, not just claimed on paper? They find only five uses (things like optical earth imaging, detecting methane leaks, eavesdropping on radio signals, and relaying broadband internet) have genuinely been demonstrated in flight — and even those rest on just four actual flight programs. It's a useful gut-check for investors, governments, and engineers deciding whether to bet on this 'high-altitude economy.'

Technical view

The authors apply a strict evidence rule — counting a function as flight-validated only given operationally relevant data return at ≥18 km — across nineteen candidate HAPS functions spanning sensing, navigation, and communication, benchmarked against existing satellite/UAV implementations. Only five functions (optical Earth observation, hyperspectral imaging, methane imaging, RF/SIGINT, and broadband relay) meet the bar, and these derive from just four flight programs (2020-2026 window), indicating the field's architectural/market narrative significantly outpaces demonstrated capability. Useful as a citable baseline for anyone modeling HAPS market sizing, technology-readiness assessments, or investment/regulatory decisions, since it separates proven capability from aspirational roadmap claims.

arXiv · cs.ARBuildable

GoalEvolve: From Handcrafted Algorithm Priors to Goal-Driven Evolution of Physical Design Algorithms

An AI 'teacher' figures out exactly which step of chip design is the bottleneck, then fixes only that.

Designing computer chips involves a long pipeline of automated steps (like placement, routing, timing optimization), and improving one step can accidentally make a later step worse — so trial-and-error tweaking often backfires. GoalEvolve is a framework where an AI system doesn't just chase generic 'better numbers' at each stage, but instead looks at the final overall quality target, figures out precisely which requirement is furthest from being met and which specific stage in the pipeline is responsible for that shortfall, and then has an AI 'Teacher' guide a focused search for a fix at exactly that spot. This targeted, goal-driven approach avoids wasted effort optimizing things that don't actually move the needle on the final chip's quality, which matters because chip design cycles are extremely expensive and slow.

Technical view

GoalEvolve reframes program-evolution for physical-design (EDA) algorithms around end-to-end QoR (quality of results) rather than stage-local objectives: it converts multi-objective target gaps into normalized deltas, identifies the dominant bottleneck via stage-resolved checkpoint evidence, and uses an LLM-based Teacher to narrow the search to the responsible algorithmic decision/source region before dispatching parallel evolutionary search there. This directly addresses the credit-assignment problem in multi-stage EDA flows where local improvement metrics are only weakly correlated with final QoR. Practitioners building EDA autotuning or LLM-guided algorithm-search pipelines could reuse the checkpoint-based bottleneck localization and goal-gap normalization as a general technique beyond physical design.

arXiv · cs.DBBuildable

Bounded Semantic Planning and Deterministic Compilation for Reliable Enterprise Text-to-SQL

Splitting 'understand the question' from 'write the SQL' turns AI database queries from unreliable to nearly perfect.

When you ask an AI to translate a plain-English question directly into a database query (SQL), it has to simultaneously figure out what you mean AND write correct code — and it can write code that runs fine but quietly answers the wrong question, like joining the wrong tables or averaging when it should be summing. This paper's system splits that job: an AI 'planner' has a back-and-forth conversation to pin down exactly what relationships and calculations the question requires, choosing only from pre-approved, schema-specific options, while a separate, non-AI program then mechanically builds the actual SQL and double-checks it. On a tricky 38-question business benchmark, this approach was judged fully correct 97% of the time versus 55% for asking the AI to write SQL directly — a huge reliability jump. This matters a lot for using AI safely in enterprise settings where a wrong-but-plausible-looking database answer could drive a bad business decision.

Technical view

Semantic path compilation (SPC) moves the LLM's role to a bounded, multi-turn planning stage that grounds phrases and selects from question-specific governed options (relationship roles, aggregation grain), while graph traversal, role predicates, grain lowering, SQL construction, and deterministic checks are handled entirely by deterministic code — narrowing where stochastic error can enter. On a 38-question adjudicated ACME insurance benchmark (3 runs/question), SPC achieved 97.4% full-run correctness (37/38) vs. 55.3% (21/38) for direct DDL-to-SQL generation, with paired discordance strongly favoring SPC (16-0). This is a concrete architecture pattern — constrain LLM output to a governed option space, push execution/validation into code — that enterprise text-to-SQL builders could adopt directly to reduce silent semantic errors (wrong join role/grain) that don't manifest as SQL execution failures.

arXiv · cs.AIBuildable

Cost Scales with Change, Not Corpus Size: Incrementally Maintaining an Evolving Semantic Substrate

Updating a document search index should cost you for what changed, not for everything you already indexed.

AI search and question-answering systems often build a fresh 'understanding' of an entire document collection every time you ask a question, which is wasteful — like re-reading your whole library every time someone asks a question about one new book. This paper proposes instead compiling the meaning of documents into a compact, reusable index once, when each document is added, and then just updating that index incrementally as new documents come in, rather than rebuilding it from scratch. The worry has always been that the math involved (a technique called SVD, essentially a way of compressing information to find its core patterns) is too expensive to update incrementally, and that switching to a better AI embedding model would force starting over completely. The authors show, on a controlled test growing from 3,000 to 9,000 documents, that the cost of updating actually depends on how much content changed, not on how large the whole collection has grown — meaning large, evolving knowledge bases can be kept up-to-date cheaply.

Technical view

The paper argues for treating semantic indexing as a compiler (build once at ingest, consult repeatedly) rather than an interpreter (re-derive meaning per query), addressing the specific bottleneck of incrementally maintaining a truncated SVD-based semantic substrate under both document churn and embedding-model changes. On a controlled synthetic pilot (dimension 256, rank 32, corpus grown 3,000→9,000 documents over 50 update events), they empirically demonstrate maintenance cost scales with the volume of change rather than total corpus size, implying incremental SVD updates (rather than full recomputation) are tractable at scale. This is directly actionable for RAG/agentic QA system builders maintaining large, frequently-updated corpora, suggesting incremental low-rank update algorithms as a viable alternative to periodic full re-embedding/re-indexing pipelines.

arXiv · cs.NIRunnable

OVS Meets PQ-TLS: Exploring Post-Quantum TLS for SDN's Southbound API

Testing whether networks can survive quantum-computer-proof encryption without slowing to a crawl.

Software-defined networking lets one central controller manage a whole fleet of network switches, and it talks to those switches over an encrypted channel called TLS, currently secured with math that a powerful enough quantum computer could someday break. This paper swaps in newer 'post-quantum' encryption methods designed to resist quantum attacks, either alone or mixed with the old methods as a safety net. The researchers actually build this into a real SDN system and measure how much slower connections get and how much extra computer processing power it burns. The point is to find out, before quantum computers are a real threat, whether we can upgrade critical network infrastructure without wrecking its performance.

Technical view

The authors implement post-quantum TLS (PQ-TLS) for the SDN southbound API (e.g., OpenFlow's TLS channel) linking controllers to switches, benchmarking pure and hybrid PQ modes (combining classical and PQ signatures/KEMs) against legacy RSA/ECDSA TLS. They measure handshake latency and CPU utilization across multiple NIST security levels and compare different PQ signature and key-establishment algorithm choices (e.g., ML-KEM, ML-DSA style schemes). The proof-of-concept quantifies the practical overhead of PQC migration in a latency-sensitive control-plane setting, giving network operators concrete data for choosing algorithms and hybrid configurations ahead of a quantum-safe transition.

arXiv · cs.LGBuildable

One Residual with Three Reuses: A Wristband Front End for Gesture Sensing

One tiny chip trick lets a smartwatch-sized sensor read hand gestures all day on a coin battery.

Wearable gesture sensors — think a wristband that reads hand movements for controlling devices or tracking tremor symptoms — usually need two very different sensors, a motion chip (like in a smartwatch) and a tiny radar chip, each running its own always-on power-hungry logic. This paper designs one shared piece of circuitry that does triple duty: deciding when to 'wake up' the classifier, choosing whether to trust the radar or the motion sensor at any moment, and improving the math that fuses their signals — all reusing the exact same small chunk of chip logic instead of three separate ones. It's tiny (a few kilobytes of memory) and runs on a real low-power microcontroller, and they tested it across four public gesture datasets. The payoff is a wristband that can sense gestures continuously for things like Parkinson's monitoring without draining a coin-cell battery.

Technical view

The authors present a wristband sensor front end that fuses a MEMS IMU with a 60GHz FMCW radar via a single shared 'residual generator' circuit, reused for three functions: classifier wake-gating, mmWave-vs-IMU sensor routing, and innovation-based EKF measurement reweighting. The shared block is extremely lightweight — 14.4KB program memory, 278B state, 110K MACs/frame — and runs on an Ambiq Apollo4 Blue Plus edge MCU, validated against four public datasets (IPN Hand, SHREC 2021, MiliPoint, EAT-Radar). This is a concrete template for resource-constrained always-on sensor fusion: reusing one residual-computation primitive across gating, routing, and filtering cuts silicon/memory footprint versus separate subsystems, directly applicable to other multi-modal edge-sensing designs under strict power budgets.

arXiv · cs.CRBuildable

LLMs for Zero-Shot Threat Detection via Structured Risk Indicators

An AI reads employee activity logs like a detective's timeline to spot insider threats it's never been trained on.

Companies worry about insider threats (employees misusing access) and stealthy long-term hacks called APTs, but training a detector for these is hard because real examples are rare and every organization's normal behavior looks different. This paper uses a large language model — the same kind of AI behind chatbots — to read a person's security logs as a timeline of events, and gives it that person's own past behavior as context, like a detective referencing a suspect's history. Instead of jumping straight to a verdict, it first writes out specific 'risk indicators' in plain language (things a human analyst would flag), then reasons over indicators across multiple time windows to catch patterns that unfold gradually. It's 'zero-shot' meaning it needs no threat-specific training data, and it outperformed simpler approaches on two established test datasets, which matters because it could make threat detection deployable without needing a mountain of labeled attack examples.

Technical view

The framework runs two LLM stages over heterogeneous security logs modeled as per-user chronological timelines, using retrieval-augmented generation (RAG) to inject each user's historical behavioral context: stage one extracts structured, interpretable threat-specific risk indicators rather than classifying raw logs directly; stage two jointly classifies across temporal sequences of these indicators to capture attack patterns spanning multiple time windows. Evaluated zero-shot on CERT r5.2 (insider threat) and PicoDomain (APT detection) across four open-weight LLM/retrieval configurations, all configurations reportedly beat baselines. The structured-indicator intermediate representation makes the pipeline auditable and reusable — practitioners can swap in different LLMs, tune the indicator schema, or extend the RAG context store for other log-based anomaly detection tasks without retraining a classifier.

arXiv · cs.DCConceptual

Vantage: Availability-Graded Broadcast for Signature-Free BFT

A blockchain-style protocol proves who-said-what fast, without expensive digital signatures.

Distributed systems like blockchains need many computers to agree on an order of events even if some participants are faulty or lying, and normally that requires digital signatures — cryptographic proof that lets one party show a third party 'X really said this.' Signatures are expensive, so 'signature-free' protocols skip them but pay a different cost: they must fully confirm every single data block is available before it can be added to the order, which is slow. This paper introduces a new building block that lets the system mark most data as 'good to go' quickly based on direct confirmations, while only the leftover uncertain bits get resolved later and more carefully — like triaging urgent cases first instead of processing everything with equal rigor. The result is a faster way to reach agreement across many computers using only cheap authentication and hashing, no costly signatures required.

Technical view

Vantage is a partially synchronous BFT protocol for n≥3f+1 parties relying solely on authenticated channels and collision-resistant hashing (no digital signatures), where parties publish blocks on hash-linked per-author lanes and each view's proposer combines a quorum-verified 'core manifest' of lane frontiers with an optimistic 'tip manifest' of freshly received blocks. The key primitive, Availability-Graded Broadcast (AGB), makes the core portion irrevocable once a quorum of first-hand responses confirms it, while grading (rather than blocking on) the unresolved tip, which gets sealed later. This decouples ordering throughput from full-block availability confirmation, letting practitioners build high-throughput signature-free consensus systems that reduce the classic latency/CPU cost of per-block availability votes seen in prior signature-free BFT designs.

arXiv · cs.DBBuildable

FROG: Efficient Range-Filtering Approximate Nearest Neighbor Search on GPUs

A GPU-native search index finds 'similar items within a price range' at massive speed.

Vector databases power things like image or product search by finding items 'similar' to a query, but often you also want to filter by a numeric attribute — like 'similar shoes under $50.' Doing both efficiently is hard, and existing methods weren't built to take advantage of GPUs, which are great at massive parallel work but need data organized just right. This paper designs an index specifically for GPUs that avoids stitching together lots of small locally-optimized pieces (which forces slow, winding searches) and instead builds one globally-aware structure so the GPU can search efficiently no matter how selective the numeric filter is. It matters for any app doing large-scale similarity search with filters — recommendation systems, image search, RAG pipelines — where speed and scale are critical.

Technical view

FROG is a GPU-oriented index for range-filtering approximate nearest neighbor search (RFANNS) that replaces the common approach of building multiple locally-optimal subgraphs (which cause long search trajectories and redundant distance computations on GPU) with a single globally-aware, vertex-centric graph structure. Each vertex organizes diverse expansion-neighbor candidates in a GPU-friendly layout, aiming to keep search efficient across the full range of filter selectivity rather than degrading on highly selective range predicates like generic GPU filtering or CPU-based indexes do. This targets a concrete systems gap — practitioners building high-throughput filtered vector search (recommendation, RAG retrieval) can adopt this index design to get selectivity-robust, low-latency k-ANN on GPU hardware.

arXiv · cs.DBBuildable

Efficient Privacy-Preserving Range Filtered Approximate Nearest Neighbor Search

Searching encrypted databases for 'similar items in a range' without the cloud ever seeing your data.

When companies store their vector databases (used for similarity search) on cloud servers they don't fully trust, they'd like the cloud to run searches — including numeric range filters, like 'similar and priced between $10-$20' — without ever seeing the actual data or queries in plain form. This is the first paper to seriously tackle that combined problem: filtered similarity search over encrypted data. Their trick is splitting the work — the user's own device figures out which parts of a pre-built index tree are relevant to the requested range and only asks the untrusted server to do the heavy similarity search on those specific encrypted parts, avoiding constantly running expensive encryption operations on everything. This matters for any business that wants cloud-scale search convenience without handing sensitive data to the cloud provider in the clear.

Technical view

The paper formalizes and solves privacy-preserving range-filtered ANN search (RFANNS) over outsourced, encrypted vector databases against an honest-but-curious server, protecting vectors, attributes, and queries. Its core technique separates range localization from encrypted vector search: the client maps a query's numeric range to a compact set of nodes in a local N-ary attribute tree (kept client-side), then the server performs proximity-graph search restricted to only the corresponding encrypted sub-indices, minimizing costly cryptographic operations applied per query. This gives a concrete architecture — client-side range-to-node mapping plus server-side restricted encrypted graph search — that implementers of secure outsourced vector databases can adopt to bound both privacy leakage and computational overhead.

arXiv · cs.ARConceptual

TRACE: Traversal and Reasoning Algebraic Computing Engine for Formal Hardware Verification

A new engine speeds up mathematically proving that chip circuits like multipliers actually compute correctly.

Before a computer chip ships, engineers need to formally prove that its arithmetic circuits — adders, multipliers, and the combined multiply-add units everywhere in AI chips — actually do the math right, and one powerful way to do this treats the circuit's logic as polynomial equations to be verified with computer algebra. But as chips get more complex, this verification can become a huge computational bottleneck, chewing through memory and time. This paper introduces TRACE, a framework that lets researchers experiment with different strategies for traversing and simplifying these polynomial proofs, and unlike older tools that mainly handled multipliers, it flexibly supports adders and multiply-accumulate units too. The goal is faster, more memory-efficient formal proofs so chipmakers can verify increasingly complex AI-era hardware without the process becoming impossibly slow.

Technical view

TRACE is a Symbolic Computer Algebra (SCA) framework for formal verification of arithmetic circuits (adders, multipliers, MAC units), representing pseudo-boolean circuit functions as polynomials and studying how different traversal/reduction strategies affect proof efficiency, memory usage, and verification time. It generalizes beyond prior SCA tools that focused mainly on multipliers, offering a flexible testbed for comparing strategies across a broader class of arithmetic primitives central to AI accelerator datapaths. Practitioners in formal EDA verification could use TRACE to benchmark or extend traversal heuristics for polynomial-based equivalence checking, targeting the memory/runtime bottlenecks that limit SCA scalability on modern high-complexity circuits.

arXiv · cs.DCBuildable

GPU implementation of a resource-constrained virtual machine

A tiny 'frugal' virtual computer gets ported to run — and race — on your graphics card.

Uxn is a deliberately minimal virtual machine (a fake computer that runs simple programs) designed to work forever on cheap, low-power hardware instead of demanding ever more resources the way modern software does. The researchers took this tiny system and ran it on a GPU, the chip normally used for graphics and AI math because it can do thousands of things at once. The trick was rewriting Uxn's assembly-like language so programs could be split into many parallel pieces, similar to how multi-core CPU code gets parallelized with OpenMP. The payoff: even a modest built-in GPU could run demanding tasks up to 19 times faster, showing 'small and simple' software doesn't have to mean 'slow.'

Technical view

The authors implement Uxn, a resource-constrained stack-based VM, as a GPU target and introduce an OpenMP-style parallel API for Uxntal to expose data parallelism across GPU threads. They evaluate on integrated GPUs, showing a 19x speedup on the compute-bound Stencil benchmark and improved frame rate on the graphics-bound Bunnymark benchmark versus baseline execution. The core contribution is a parallelism abstraction letting frugal, minimal-ISA VM code exploit SIMT hardware without abandoning the VM's resource-constrained design philosophy. This offers a template for retrofitting other constrained/embedded VMs with GPU-parallel execution paths.

arXiv · cs.DCBuildable

MELD: A Protocol for Merging Knowledge Across Distributed Agentic Memories

A protocol lets AI agents' separate memories merge facts, spot contradictions, and stay in sync.

When multiple AI agents each keep their own memory of what they've learned, there's currently no good way for them to compare notes — one might phrase a fact differently than another, or two agents might flatly disagree, and today's systems just let one version silently overwrite the other. MELD is a proposed protocol that treats each agent's memory as a live knowledge graph (a network of connected facts) and, whenever new information arrives, decides whether to add it, merge it with something similar, link it to related facts, flag it as a conflict, or reject it. It figures this out using three checks — do the facts share an identity, are they semantically similar, and does a language model judge them as agreeing or contradicting — and every change happens through a single auditable 'patch' so nothing gets lost quietly. The goal is trustworthy, decentralized collaboration between AI agents that each keep control of their own memory while still reconciling what they collectively know.

Technical view

MELD is a coherence protocol for federated agentic memory that models each agent's memory as a knowledge graph and admits incoming claims via a five-outcome decision procedure (insert, merge, relate, conflict, reject). Decisions combine three signals — scoped claim-key identity matching, embedding similarity, and NLI (natural language inference)-based entailment/contradiction verdicts — gated by context and freshness checks, with state mutation restricted to a single auditable, authenticated Patch object. It binds to standard pub/sub transports using a per-claim status CRDT (conflict-free replicated data type) to keep memory ownership sovereign per agent while allowing eventual convergence across the federation. Practitioners building multi-agent systems could adopt this as a reconciliation layer instead of ad hoc memory-sharing or naive overwrite semantics.

arXiv · cs.ARBuildable

Beyond Binary Priorities: Multi-Tier SLA Scheduling for Large Language Model Serving

LLM servers learn to juggle many customer priority tiers, not just 'important' vs 'everyone else.'

When companies run large language models as a service, different customers have different needs — some need instant replies, others are fine waiting for batch jobs overnight — and the scheduling software deciding whose request runs next matters a lot. An existing system called Llumnix could only tell requests apart as 'high priority' or 'normal,' too crude for businesses selling multiple pricing tiers. This paper extends that scheduler to handle any number of priority tiers and tests it against three realistic customer mixes (evenly spread, bell-curve, and typical enterprise patterns) using a detailed LLM traffic simulator. The point is making sure paying customers actually get the service level they're promised without wasting computing capacity.

Technical view

The work generalizes Llumnix's binary (high/normal) priority scheduling for LLM inference serving to an arbitrary number of SLA tiers, implementing per-tier headroom within its migration-capable, freeness-metric-based multi-instance scheduler. Evaluation uses Vidur, a high-fidelity LLM inference simulator, across uniform, Gaussian, and enterprise-shaped priority distributions to measure how well tier-differentiated SLOs are met under load balancing, defragmentation, and auto-scaling. This targets production deployments needing finer-grained QoS differentiation than binary priority allows, and the per-tier headroom mechanism is a concrete extension point for practitioners running Llumnix-style schedulers. Replicable via the open Vidur simulator and Llumnix codebase.

arXiv · cs.NIConceptual

Towards the Interplanetary Internet: An IoT Perspective

Building an internet for Mars and the Moon may reuse the same tricks as your smart thermostat.

As space agencies and companies plan permanent robotic and human outposts on the Moon and Mars, they need networking that can handle huge delays and unreliable links across millions of miles — something ordinary internet protocols weren't built for. This paper argues that instead of inventing entirely new deep-space networking from scratch, engineers should reconsider standard Internet Protocol (IP), because deep-space communication has a lot in common with 'Internet of Things' (IoT) scenarios — think low-power sensors with spotty, delayed connections here on Earth. It walks through those similarities, summarizes relevant international standards work already underway, and lays out research directions for adapting IoT-style IP protocols to work between planets. The bigger idea is avoiding reinvention of space networking by borrowing lessons already learned connecting billions of small, constrained devices on Earth.

Technical view

This is a position/survey paper arguing for adopting IP-based IoT protocol stacks for the Interplanetary Internet, drawing analogies between deep-space link constraints (long delay, intermittent connectivity, constrained nodes) and terrestrial IoT scenarios. It surveys relevant IETF standardization efforts and identifies where existing IoT protocols (e.g., constrained routing, delay-tolerant mechanisms) could be reused or adapted rather than building bespoke deep-space protocols. The contribution is conceptual framing plus a research agenda rather than new experimental results. Useful as a starting map for networking researchers or standards contributors deciding where to focus effort in space-IoT protocol design.

arXiv · cs.DCBuildable

DB-SpMSpV: Dual-View Blocked Sparse Matrix-Sparse Vector Multiplication for Dynamic GPU Workloads

A smarter way to multiply 'mostly-empty' matrices on GPUs adapts on the fly to how empty things really are.

Multiplying sparse matrices (grids of numbers where almost everything is zero) by sparse vectors is a core operation behind graph algorithms, recommendation systems, and AI inference, but making it fast on a GPU is tricky because how sparse the data is can change from one moment to the next. Existing approaches lock in one storage format and traversal strategy ahead of time, which works well for some sparsity patterns but poorly for others. DB-SpMSpV instead chops the matrix into small fixed-size blocks, keeps two different 'views' of the same data ready to go, and at runtime picks whichever traversal path and low-level routine best matches the current sparsity pattern — without storing the data twice or adding scheduling overhead. This adaptive flexibility means the same code can stay fast across the unpredictable, shifting sparsity patterns real workloads actually produce.

Technical view

DB-SpMSpV partitions sparse matrices into fixed-size 2D blocks and maintains dual block-level CSR/CSC (compressed sparse row/column) views over a single shared low-level block payload, enabling both row-driven pull and column-driven push traversal without duplicating storage. At runtime it selects the global traversal path based on input block sparsity and picks per-block microkernels accordingly, decoupling storage layout, traversal direction, and kernel choice — a departure from prior GPU SpMSpV designs that bind these together statically. This targets graph traversal, sparse linear algebra, and sparse model inference workloads with dynamically varying vector sparsity. Practitioners building sparse GPU kernels can adopt the dual-view block layout as a drop-in mechanism for runtime traversal-path adaptation without extra memory overhead.

arXiv · cs.DCBuildable

DepTGL: A Parallel Framework for Memory-based TGNN Training with Adaptive Temporal Data Dependency Management

Training AI models that remember graph history over time gets faster by rethinking who waits for whom.

Some graph neural networks (AI models that learn from networks of connected data, like social graphs or transaction networks) keep an evolving memory of each node's past behavior, updated as new events stream in — these are called memory-based temporal graph neural networks. Training these across multiple machines is hard because updates must happen in strict time order, forcing machines to constantly wait on each other and creating traffic jams, especially when events aren't spread evenly over time. DepTGL redesigns how this data dependency is tracked and managed, mixing local caching of recent events with selectively sending only the updates that truly matter to other machines, plus a smarter way to keep cached copies synchronized during learning. The result is training that scales better across machines by cutting wasted waiting and communication.

Technical view

DepTGL is a distributed training framework for memory-based Temporal Graph Neural Networks (M-TGNNs) that restructures temporal-dependency management from a data-centric angle rather than enforcing strict global chronological ordering. It introduces a hybrid dependency scheme combining temporal-event caching with selective, dependency-driven remote communication to balance caching and communication overhead, plus a gradient-aware cache-synchronization mechanism. This directly targets the load-imbalance and synchronization-overhead problems that arise in distributed M-TGNN training under skewed temporal event streams. It's relevant to practitioners scaling temporal-memory GNN training (e.g., TGN-style models) across multiple GPUs/nodes who currently bottleneck on strict-order synchronization.

arXiv · math.OCConceptual

Stochastic Gradient Tracking over Time-Varying Networks: One-Step Lyapunov Analysis

A tighter math proof shows decentralized learning teams converge just as fast even with a wobbly network.

Imagine many computers (agents) each holding a piece of data, trying to jointly learn something by repeatedly averaging their estimates with neighbors over a network that can change and even become temporarily disconnected — this is decentralized learning. Prior analyses typically required either a stable, well-connected network or complicated tracking of dynamics across many communication rounds at once. This paper proves that as long as a 'window' of several consecutive communication rounds reliably shrinks disagreement between agents on average, you can build one clean mathematical tool capturing the whole system's progress step by step. Using this, they show decentralized learning converges essentially as fast as if all the data were on one central computer, once the network has run long enough — a reassuring theoretical guarantee for distributed AI training over unreliable connections.

Technical view

The paper analyzes decentralized stochastic gradient tracking over time-varying networks under a uniform window-mixing condition, where products of τ consecutive doubly stochastic mixing matrices contract disagreement by factor λ<1 even though individual matrices/graphs may be non-contracting or disconnected. The key contribution is a time-varying quadratic Lyapunov norm that converts this window-level contraction into an exact one-step Lyapunov identity, yielding coupled one-step recursions for centroid and disagreement error without unrolling dynamics over the full mixing window. They prove leading stochastic error terms of Õ(1/(NK)) for smooth strongly convex objectives and O(1/√(NK)) for smooth convex objectives, matching centralized mini-batch rates and giving linear speedup in N after a network-dependent transient. This provides sharper, more general convergence theory researchers can apply to prove rates for decentralized SGD-based algorithms under weaker, disconnection-tolerant network assumptions than prior work.

arXiv · cs.NIBuildable

SbDN: Source-based TSN-Grade Deterministic Networking using Commodity Switches

Skip pricey smart switches — do all the network traffic-cop work at the source instead.

Factories, cars, and planes need networks where messages absolutely, positively arrive on time — a robot arm command can't be a millisecond late. The usual fix, Time-Sensitive Networking, needs every switch along the path to be a fancy, expensive, carefully-configured piece of hardware. SbDN's trick is to move all the scheduling smarts into one central 'brain' made of cooperating software agents, and have it tell only the sending devices when to transmit, so ordinary cheap switches can just pass packets along without any special intelligence. This could make deterministic, guaranteed-timing networks far cheaper and easier to reconfigure on the fly.

Technical view

SbDN replaces per-hop TSN switch configuration with a centralized multi-agent controller that computes schedules and enforces them solely at source endpoints, treating commodity switches as dumb forwarders. Its core mechanism, Temporal Network Partitioning (TNP), achieves strict temporal isolation between traffic classes without switch-side gating or shaping. Practitioners could evaluate this as a lower-cost alternative to full TSN switch fabrics, particularly where topology or traffic patterns change at runtime and per-switch reconfiguration overhead is prohibitive.

arXiv · cs.OSBuildable

AdaSprite: Resource-efficient Online Co-Adaptation for V2I Systems Under Large-scale Data Drifts

Roadside AI cameras keep 'seeing' differently as traffic and lighting shift, so they quietly retrain themselves.

Smart intersections and highways increasingly use AI vision models that combine feeds from many cameras to understand traffic — think of it as many eyes reporting to one brain. The problem is that what those cameras see keeps changing over minutes to hours (weather, new vehicle types, different drivers), which confuses the AI's internal 'expert' modules over time. AdaSprite lets multiple such systems, running on nearby roadside computers rather than a distant cloud, help each other adapt to these changes efficiently, avoiding the delay and privacy risk of sending video far away while still using less computing power than each camera fixing itself alone. It's about keeping AI traffic systems accurate and fast without needing constant expensive retraining.

Technical view

The system targets Vision Mixture-of-Experts (V-MoE) backbones for VLM-based vehicle-infrastructure perception, where sparse expert routing enables conditional computation but degrades under sustained distribution shift across viewpoints. AdaSprite proposes resource-efficient online co-adaptation across multiple edge-deployed V-MoEs, coordinating updates among edge servers rather than relying on cloud retraining or isolated on-device fine-tuning. This addresses a real deployment gap for multi-camera V2I perception pipelines that must stay accurate under compute and latency constraints as scene statistics drift.

arXiv · cs.DCBuildable

Agent-Native Telemetry: Verifiable State-Delta Evidence for Autonomous Operations

Instead of writing logs for humans, give AI systems tamper-proof snapshots of exactly what changed.

When a computer system logs what it's doing, it usually writes verbose sentences meant for a person to read — but AI agents now often 'read' those logs instead, wasting effort parsing English-like text just to figure out what actually changed. Agent-native telemetry proposes recording just the state changes themselves — like a structured diff — plus cryptographic proof that nothing was missed or faked. It organizes facts into four building blocks (transitions, observations, relationships, and checkpoints) that are digitally signed and chained together. The payoff is AI systems that can trust and reason over operational history efficiently, which matters more and more as AI agents start running infrastructure autonomously.

Technical view

The Agent Telemetry Protocol (ATP) and State-Delta Evidence Ledger restructure operational logging around four evidence primitives — Transitions, Observations, Relations, and State Checkpoints — using content-addressed, cryptographically verifiable structures instead of prose log lines, aiming to cut token/context overhead for LLM-based operators while adding provenance and completeness guarantees. This is essentially a schema-and-integrity layer analogous to a signed event-sourcing ledger, purpose-built for machine consumption rather than human readability. Engineers building autonomous ops tooling could adopt ATP as a logging substrate to reduce agent context costs and add auditability without redesigning existing state machines.

SW

Software & Programming

50 new
arXiv · cs.AIConceptual★ flagship

When Agents Coordinate: Measuring Coordination in Multi-Agent AI Coding

When teams of AI coders work together, this measures how they actually talk and share files.

People are starting to run whole teams of AI coding agents on one task, but we usually only track whether they finished and what it cost — not how they actually coordinated inside. This work builds a measuring instrument that turns each run into a time-stamped network: agents and files are dots, and messages, file writes, and file reads are the connecting arrows, each with a cost. They ran this over 1,902 runs while varying team size, org structure, and rules about who can touch which files, then watched how coordination patterns shift as teams grow. One striking finding: direct messaging balloons almost with the square of team size, largely from an initial round of everyone introducing themselves. It matters because understanding coordination overhead is key to designing multi-agent systems that scale instead of drowning in chatter.

Technical view

The paper introduces a temporal-network instrument for multi-agent coding: each run is a graph with agents and files as nodes and timestamped, directed, cost-weighted edges for messages, writes, and reads. They apply it to 1,902 test-suite-evaluated runs across a factorial of team size, team structure, and file-access policy, yielding empirical scaling laws for coordination. A headline result is that direct messaging grows roughly quadratically with agent count, much of it an early 'introductions' burst. Practitioners can adopt the network representation as an evaluation layer beyond pass/cost to diagnose communication overhead, compare topologies and file policies, and design multi-agent configurations that curb superlinear messaging growth.

arXiv · cs.SERunnable★ flagship

TDD-Agent: Test-Driven Reasoning for Code Generation

Make an AI write the tests first, then code to pass them — like disciplined human developers.

Language models write code well but stumble on big, real-world codebases where correctness is hard to guarantee. A common patch is to generate tests and check the code against them afterward, but if those tests are themselves wrong or incomplete they give misleading feedback. TDD-Agent instead borrows test-driven development from human engineering: it first has the model write executable tests to pin down the expected behavior before writing any implementation, which forces it to clarify what 'correct' even means. Then it refines both the code and the tests together in a loop, using what actually runs and fails as feedback for each. It matters because test-first reasoning consistently boosts correctness, and treating tests as an evolving partner rather than a fixed judge avoids being misled by flawed tests.

Technical view

TDD-Agent operationalizes test-driven development for LLM code generation. It first prompts the model to generate executable tests to elicit behavioral specifications, then performs iterative dual-track refinement over both code and tests using execution feedback, rather than treating generated tests as static post-hoc validators. An isolated prompt variant, TDD-prompt, shows test-first reasoning alone consistently improves over standard reasoning on LiveCodeBench, and the full agent extends this to repository-level tasks. Practitioners can adopt the test-first prompting and the co-evolving code/test refinement loop to reduce reliance on possibly-incorrect static test oracles and improve correctness on complex, multi-file generation.

arXiv · cs.ARBuildable

NeuroAbs: A Neuro-Symbolic RTL Abstraction Framework for Property Checking Acceleration

Using AI to figure out which chip wiring can be safely simplified before a computer proves it correct.

Verifying that a hardware chip design (RTL code) works exactly as intended is a slow, brute-force process, especially as chips get more complex. Engineers use 'abstraction' — temporarily hiding or simplifying parts of the design — to speed up the proof, but deciding what's safe to simplify has traditionally needed either manual expert effort or rigid rule-based tools. NeuroAbs uses a large language model (LLM) to read the chip's code and intelligently spot which signals are good candidates for abstraction, then combines that AI judgment with a structured, symbolic representation of the code to actually carry out the simplification accurately. Each simplification is checked afterward to make sure it doesn't change what's being verified, blending flexible AI insight with rigorous correctness guarantees.

Technical view

NeuroAbs is a neuro-symbolic RTL abstraction pipeline for accelerating formal property checking: an LLM performs RTL analysis to propose candidate signals for abstraction, and a second LLM stage generates the abstraction transformation guided by an AST-based symbolic representation of the RTL, aligning the LLM's edits with the intended structural transform rather than free-form text generation. Each abstraction is soundness-checked before being applied to the property-checking flow, aiming to combine LLM flexibility with formal guarantees absent from prior rule-based abstraction tools. This is directly usable by hardware verification teams as a preprocessing step ahead of existing model checkers to reduce state-space size on complex designs.

arXiv · cs.LOConceptual

Approximate Functional Dependencies---Implication Problem Revisited

Databases with a few sloppy rows still need rules — this paper finds gaps in how we reason about 'mostly true' rules.

A functional dependency is a rule like 'if two records share the same ID, they must share the same name' — a way of keeping databases consistent. Real-world data is messy, so researchers study a relaxed version that tolerates a small number of rule-breaking rows, and prior work wrote down logical rules for reasoning about combinations of these relaxed dependencies. This paper shows those earlier rules were incomplete — there's a valid logical consequence they miss — and pins down exactly where the reasoning still works (for a simpler 'single-column' case) and studies how hard it is to check whether a given approximate rule holds. It matters because these dependency rules power query optimization and data-cleaning tools, so getting the underlying logic right keeps those tools trustworthy.

Technical view

The paper revisits the axiomatization of approximate functional dependencies (a quantitative relaxation of classical FDs tied to dependence atoms in team semantics/team logic, per Väänänen 2017), showing the existing inference rule system is incomplete — it fails to derive a semantic consequence that does hold. The authors prove the axiomatization is complete when restricted to unary dependencies, and analyze the computational complexity of model checking for approximate dependence atoms. This tightens the theoretical foundation underlying approximate-FD-based data cleaning, discovery, and constraint-based query optimization, and flags open completeness questions for the general multi-attribute case.

arXiv · cs.SEConceptual

Revisiting the Performance of Generative Artificial Intelligence on Introductory Object-Oriented Programming Assessments: Insights from 2026

Five top AI chatbots took a real coding class exam — and beat the average student.

This study asks a very practical question: if you gave today's best AI chatbots the same programming tests and exams used in a real university's introductory object-oriented programming course, how would they score? The researchers ran five well-known AI systems — including recent versions of ChatGPT, DeepSeek, Gemini, Claude, and Microsoft Copilot — through the same assignments, graded them exactly like student work, then compared results against actual historical student performance and last year's version of this same study. Every AI system outperformed the average student, and the researchers also catalogued the specific mistakes these models still make. This matters because it's hard evidence for how educators should rethink homework, exams, and academic integrity as AI coding ability keeps climbing year over year.

Technical view

The study benchmarks five contemporary GenAI systems (ChatGPT-5.2, DeepSeek-V3, Gemini 2.5 Flash, Claude Sonnet 4.5, M365 Copilot) on authentic introductory OOP course assessments, grading outputs with the same rubrics used for students and comparing against historical cohort performance and a prior-year study for longitudinal comparison. All five models exceeded the average student score, and the authors perform an error analysis characterizing recurring model failure modes on OOP-specific tasks. This provides a reusable evaluation methodology (real course assessments plus student-baseline comparison and error taxonomy) that educators or researchers could replicate on their own curricula or newer model releases. It's most useful as empirical grounding for CS-education policy decisions around AI-assisted assessment integrity.

arXiv · cs.LOConceptual

A Kernel-Checked Exclusion Certificate for Erdős Problem 647

A math proof about number-of-divisors patterns, checked by a computer so rigorously nothing is taken on faith.

Erdős left behind an open question about a function that counts how many divisors a number has (τ), asking whether some pattern in its growth ever breaks past a certain point. Earlier computer searches checked huge ranges of numbers for counterexamples, but those searches relied on 'trust me, the code ran correctly' rather than a fully verified logical proof. This paper redoes the check for numbers up to a billion using Lean, a proof assistant that verifies every logical step from rock-bottom axioms, with no shortcuts or unproven computational leaps. The result is a proof you can trust completely, not just a program that probably worked — an important distinction as math increasingly leans on computers.

Technical view

Erdős Problem 647 concerns whether max_{m<n}(m+τ(m)) ≤ n+2 fails for any n > 24; prior work excluded solutions computationally up to 10^12 and, with a Lean component using native_decide (outside the trusted kernel), up to ~9.17×10^18. This paper delivers a fully kernel-checked exclusion for 24 < n ≤ 10^9 in Lean 4 with axiom closure limited to {propext, Classical.choice, Quot.sound} — no sorry, no native_decide — via 6,685,922 concatenated factorization-witness intervals requiring only primes below 1024. It's a template for converting native_decide-backed computational number theory results into fully trustworthy, kernel-verified proofs, useful to anyone doing large-scale formalized computational mathematics in Lean.

arXiv · math.LOConceptual

Idealizing Useful Fictions in Omega Grounded Arithmetic

A logic system that lets computers say 'this will definitely finish' about their own infinite searches.

Some formal logic systems only let you claim something is true if a computation actually finished proving it — which cleverly sidesteps classic paradoxes like 'this sentence is false,' since unfinished computations just stay in limbo rather than breaking everything. But there's a gap: a system called RGA can talk about its own computations yet can't always certify that an endless search will actually terminate with a real answer. This paper adds one new rule — if you can show every individual case of a 'for all' statement is settled, then the whole statement counts as settled — and studies the resulting stronger system, OGA, with every claim double-checked by machine. It's foundational work on how much a self-referential, computation-grounded logic can safely know about itself.

Technical view

Grounded arithmetic ties assertability to terminating computations, making the logic paracomplete so undecided sentences neither assert nor refute (avoiding Liar-style explosions); RGA extends this reflectively but can't internally certify termination of its own unbounded searches. This paper adds ATI, an ω-grounded universal-closure rule (if every numeric instance is certified decided, so is the universal), yielding OGA — syntactically identical to RGA but semantically stronger — with all results machine-checked. This is relevant to researchers working on paraconsistent/paracomplete logics, reflective type theories, or formalized proof-theoretic strength comparisons between grounded systems.

arXiv · cs.SEConceptual

Reshaping the SDLC for Data- and AI-Centric Systems

Software development rules built for hand-written code don't fit AI systems whose behavior comes from data too.

Traditional software engineering assumes that if you write and test the code carefully, you know how the program will behave. But AI-powered systems' behavior also depends heavily on the data they were trained and fed with, which can silently drift away from what worked during development, breaking the old assumption. This paper reviews and organizes practices — dubbed DataOps, MLOps, and LLMOps — that treat data and models as first-class citizens alongside code throughout a project's life, from initial requirements through architecture and ongoing operation. It's a roadmap for how teams should rethink their whole development process now that 'correct code' no longer guarantees 'correct behavior.'

Technical view

The paper synthesizes software engineering, data management, ML systems, and HCI literature into a phase-structured account of how DataOps/MLOps/LLMOps practices transform the SDLC across requirements, architecture, development, and beyond for data- and AI-centric systems, motivated by the fact that behavior emerges from code+data+model interaction and degrades under real-world drift. It offers four contributions (the literature synthesis being the first named one) intended as a reference framework for teams designing lifecycle processes around continuous data/model monitoring rather than one-time code verification. Useful for engineering leads defining or auditing MLOps/LLMOps practices against a structured, research-backed lifecycle model.

arXiv · cs.SEBuildable

SpecTrum: Specification-Guided Differential Fuzzing for Ethereum Consensus Clients

Fuzzing bots stress-test Ethereum's rulebook implementations to catch bugs before they fork the blockchain.

Ethereum only stays trustworthy if all the different software 'clients' that run it agree on every single state change — if they disagree due to a bug, the blockchain can literally split in two, or worse. There's an official Python reference and a hand-written test suite meant to define correct behavior, but because the reference is just runnable code rather than a precise rulebook, nobody can be sure every edge case is actually tested. SpecTrum turns that reference into an explicit, structured specification with clear if-then conditions, then automatically generates and fuzzes test inputs against it to find gaps. It's essentially building a more rigorous, automatically-checked rulebook to prevent catastrophic blockchain bugs.

Technical view

SpecTrum targets Ethereum consensus client divergence risk by first producing Consensus-SpecTec, a mechanized specification derived from the Python consensus-spec that makes implicit runtime validity conditions explicit as if-premises, replacing the ad hoc hand-crafted spectests suite. It then uses this explicit specification to drive differential fuzzing — generating inputs that exercise validity conditions systematically — to surface discrepancies between client implementations and the reference. This gives consensus client teams (Prysm, Lighthouse, Geth-adjacent, etc.) a more systematic coverage tool for validity-condition testing than manual test-suite authoring provides.

arXiv · cs.SERunnable

What Aggregate Scores Miss: Measuring Item-Level Regressions in Commercial LLM API Migrations

A single 'GPT got better' score can hide thousands of cases where it actually got worse.

When companies upgrade the AI model behind their product, they usually check a single overall benchmark score to decide if the new version is better. This study shows that a good average can hide a lot: on nearly a thousand individual test questions, run 50 times each to see how consistent the answers are, they find that many specific questions actually got worse even as the overall number improved. Using statistical methods to weed out noise, they classify each question as truly improved, truly regressed, unchanged, or unclear. The takeaway is that teams migrating to a new AI model version should check specific use cases, not just the headline benchmark number, or they risk silently breaking things that used to work.

Technical view

The authors evaluate three pairwise GPT-5.4→5.6 upgrades across 900 public benchmark items (grad-level knowledge, olympiad math, instruction following), sampling each item 50 times per model to estimate per-item reliability, then classify items as reliably improved/regressed/equivalent/inconclusive using false-discovery-rate correction, a practical-significance threshold, and a label-permutation null for calibration. Across all nine benchmark×migration cells they find nontrivial rates of reliable item-level regression coexisting with net-positive aggregate scores. This gives practitioners a concrete methodology — repeated per-item sampling plus FDR-controlled significance testing — for auditing LLM API migrations beyond trusting vendor-reported aggregate benchmarks.

arXiv · cs.SEBuildable

GADR: Gathering Architecture Decision Records from Meeting Transcriptions

AI eavesdrops on messy meetings and writes down the tech decisions everyone forgot to document.

Software teams are supposed to keep 'Architecture Decision Records' (ADRs) — short documents explaining why they chose one technical approach over another — but in reality those decisions get made verbally in meetings and never written down. GADR uses a team of AI agents that work together and double-check each other's output to sift through raw, messy meeting transcripts (full of tangents and half-finished thoughts) and pull out the real decisions buried inside. It then drafts a properly formatted ADR from what it found. In tests on real project meetings, senior architects and students found the AI-generated drafts captured most of the important decisions and were clear and useful, doing notably better than just asking one AI model to do it in a single pass.

Technical view

GADR is a multi-agent, self-correcting pipeline that extracts architectural decisions from raw meeting transcripts and outputs Nygard-formatted ADR drafts, in contrast to prior work that assumes clean, pre-structured input. A feasibility study (5 real transcripts, 4 senior-architect reviewers, 15 student evaluators) shows it captures most expert-identified decisions and beats a zero-shot single-pass baseline on clarity and usefulness. The multi-agent extraction/self-correction/formatting stages are a reusable template for teams wanting to bolt ADR generation onto existing meeting-transcription pipelines (e.g., Zoom/Otter integrations).

arXiv · cs.SERunnable

Benchmarking Automated Security Patch Backporting: How Far Are We?

A fair stress-test shows AI patch-porting tools aren't nearly as good as their own numbers claim.

When a security bug is fixed in the newest version of software, someone often needs to 'backport' that same fix into older versions still running in production — tedious, error-prone work that tools try to automate. Vendors of these tools report success rates above 80%, but only on their own narrow test setups (one repo, one version range). The researchers built Porting Benchmark, a much bigger and more varied test set of 1,234 real cases spanning different repositories, branches, and versions, and re-tested five leading tools under the same fair conditions. The result: once you test broadly instead of on a tool's home turf, the picture of which tools actually work well changes — a caution for anyone relying on these self-reported numbers to trust automated security patching.

Technical view

Porting Benchmark is a curated dataset of 1,234 security patch backporting cases spanning cross-version, cross-branch, and cross-repository scenarios, paired with a unified evaluation harness. The authors re-evaluate five representative tools (spanning program-analysis, LLM-prompting, and LLM-agent approaches — e.g., PortGPT, TSBPort, FixMorph) under aligned settings and find that relative rankings and absolute success rates shift substantially compared to each tool's original isolated evaluation. Anyone building or comparing backporting tools now has a reproducible, heterogeneous benchmark to test against rather than relying on single-repo self-reported numbers.

arXiv · cs.SEBuildable

Unified Message Model for Heterogeneous Serial Data Exchange Protocols

A universal blueprint lets engineers describe any device's private wiring language in one common format.

Embedded systems — think car computers or factory machines — are stitched together from many sensors and controllers that each 'talk' using their own serial communication protocol, some standardized, many custom and undocumented. That variety makes it hard to build automated tools (like code generators or protocol analyzers) that work across projects, because each protocol has to be handled by hand. This paper proposes one formal, protocol-agnostic way to describe any such message: defining the data types, the small reusable building blocks ('containers'), and how they combine into a full message. On top of the model, it also gives practical methods for actually using it in real development workflows. The payoff is that automation tooling for embedded systems can be written once and reused, instead of rebuilt for every new protocol.

Technical view

The paper defines a protocol-agnostic formal message model built from typed data elements, atomic 'container' message elements, and a composable full-message structure, aimed at giving automation toolchains a machine-processable, deterministic basis for describing serial protocols ranging from standardized to fully project-specific. Beyond the model, it introduces companion methods for practical application (likely covering tooling/co-generation workflows, per the abstract's cutoff). A practitioner building embedded protocol tooling (parsers, code generators, test harnesses) could implement this model as a schema layer analogous to Protobuf/DSLs but targeted at heterogeneous serial comms.

arXiv · cs.AIConceptual

Graph Surgery and the Do-Operator: A Precise Correspondence for Acyclic Structural Causal Models

A math proof nails down exactly when 'erasing an arrow' on a causal diagram truly matches forcing that outcome in reality.

In causal reasoning, researchers use two different ways to represent 'what if we forced this thing to happen' (like imagining forcing a light switch to 'on' rather than just observing it): one is graphical — draw the cause-and-effect diagram and delete the arrows pointing into that variable — and the other is functional — actually replace the rule governing that variable with a fixed value. Everyone assumes these two moves are 'the same thing,' but that had never been made mathematically precise, since one just returns a picture and the other returns actual rules plus the values imposed. This paper proves a theorem pinning down exactly when the two match, for systems where every variable is a deterministic result of its causes and there's no circular dependency. This matters because a lot of scientific, medical, and AI causal-inference software leans on the graphical shortcut being trustworthy, and this gives it a rigorous foundation.

Technical view

For deterministic acyclic structural causal models (SCMs) with finitely many endogenous variables, the paper proves Graph(F^ι) = Surg(Graph(F), T_ι): applying the do-operator's mechanism replacement produces exactly the dependency structure obtained by graph surgery (arrow deletion) on the corresponding graph. It further characterizes when this equality holds using the model's stated graph G even when G contains unused arrows (dependencies present in the graph but absent in the actual mechanisms). This is directly usable as a correctness foundation for formal-verification or theorem-proving tools that manipulate do-calculus expressions symbolically.

arXiv · cs.FLConceptual

A New Syntax and Semantics for Probabilistic Trace Expressions

New logic lets software monitors reason honestly about the events they missed instead of pretending they saw everything.

Runtime verification is the practice of watching a running system's stream of events to check it's obeying its rules — for example, 'the door must lock before power shuts off.' Most monitoring tools assume they see every single event, but real sensors and networks drop, delay, or hide data, so that assumption often fails in practice. This paper redesigns the rulebook (a formalism called Trace Expressions) to build uncertainty into its core logic rather than bolting probabilities onto individual technical steps as an earlier 2022 version did. The result is a monitor that can give a probability-weighted verdict — 'likely compliant' rather than a flat yes/no — even when some information is missing. This matters for monitoring real-world systems like IoT devices or distributed networks where perfect visibility is never guaranteed.

Technical view

The paper redefines Probabilistic Trace Expressions (PTE), replacing the original 2022 formulation's approach of attaching probabilities to individual syntactic transitions with a new syntax and semantics where probabilities are integrated at the level of the operational semantics itself, improving compositionality and expressiveness for monitoring under partial observability (lost, delayed, or unobservable events). This gives runtime-verification tool builders a principled way to compute probabilistic verdicts and quantify uncertainty when instrumenting distributed or IoT systems with unreliable telemetry, rather than assuming complete traces.

arXiv · cs.AIBuildable

TRUSS: Towards Task-Reliable and User-Safe Automated Agent Skill Generation

A safety inspector fact-checks AI-written 'skills' and sandbox-tests their real actions before an agent gets to use them.

AI agents can be handed 'Skills' — bundled instructions plus tools for doing a specific job — and it's tempting to have AI generate these skills automatically to save effort. The problem is that just checking whether the final output looks right doesn't tell you what actions the agent will actually take when using that skill, or whether those actions are safe. TRUSS first fact-checks the skill's claims against real evidence and screens the whole package against nine predefined safety rules; only skills that pass get run by a 'shadow' agent inside a locked-down sandbox, where every tool call it makes is monitored, policy-checked, and logged before being trusted. This matters as AI agents get handed more real-world tools and permissions, since it catches both bad information and dangerous side effects before they reach production.

Technical view

TRUSS is a two-stage evidence-guided gate for auto-generated Agent Skills: a static gate that verifies functional claims against source/domain evidence and checks the artifact against nine predefined safety properties, followed by dynamic evaluation where admitted candidates run inside a shadow agent within a Controllable Execution Environment — brokered tool calls expose requested actions to policy enforcement and are recorded for audit. This static+dynamic gating pattern is directly reusable for teams building pipelines that auto-generate or vet agent tools/skills before production rollout, giving both a documentation-accuracy check and a sandboxed behavioral audit.

arXiv · cs.SERunnable

REST API Testing with Verified LLM-Inferred Dependencies and Response-Driven Refinement

Instead of trusting an AI's guess about API rules, this tool actually runs the calls to check before writing tests.

Testing a web API (like a shopping site's backend) means calling its operations in the right order — you can't delete a user account before you've created one, for instance. Recent AI-based testers read the API's documentation and guess these ordering rules (dependencies) using an LLM, but guesses can be wrong, leading to broken or incomplete test sequences. APIPilot instead treats the AI's guesses as hypotheses: it proposes dependencies using both rule-based analysis and LLM reasoning, then actually calls the live API to verify each guess is really true before using it to build test sequences. This produces more reliable, executable test suites, which matters for automating quality assurance of the web services everything from apps to bank systems depend on.

Technical view

APIPilot derives candidate producer-consumer dependencies from OpenAPI specs using structural heuristics plus LLM-based semantic reasoning, then treats these as hypotheses and validates each via concrete API executions before admitting it into a dependency graph used for test-sequence generation — directly addressing the spurious/missed-dependency problem of prior LLM-only approaches. A practitioner could integrate this as a CI pipeline stage: point it at an OpenAPI spec and a live/staging endpoint to auto-generate execution-validated test sequences rather than trusting unverified LLM output.

arXiv · cs.AIBuildable

Agent Lightning v1.0: Towards Harnessed Agentic RL

A plug-and-play bridge lets any AI agent app get reinforcement-trained without rewriting its internal scaffolding.

AI agents run inside 'harnesses' — the surrounding software that manages their tools, memory, and step-by-step control flow. Making these agents better through reinforcement learning (RL), a training method that rewards good behavior and penalizes bad, usually means deeply rewiring the agent's code to hook into the training system. Agent Lightning instead connects any existing agent harness to an RL trainer through a lightweight stand-in that looks like a normal AI model endpoint, so the harness keeps running the agent exactly as it already does while the trainer just watches the conversations flow by and learns from them. This has already been adopted by several other training frameworks, and it matters because it makes RL-based improvement accessible to real, already-built agent products without a costly rewrite.

Technical view

Agent Lightning formalizes 'harnessed agentic RL': a disaggregated architecture where an LLM-endpoint proxy connects an arbitrary, unmodified agent harness to an RL trainer, so the harness owns the environment-interaction loop while the trainer only observes sequences of LLM request-response pairs. This introduces nontrivial engineering problems — retokenization, sample merging, advantage calculation, loss normalization, backend scheduling — that v1.0 addresses, and the paradigm has already been adopted by frameworks like verl Uni-Agent, AReaL 2.0, slime, and Polar. Practitioners with an existing agent harness can plug it into this proxy to RL-post-train the underlying model without modifying agent logic.

arXiv · cs.SEBuildable

Beyond FLOPs: Energy-Aware Knowledge Distillation for Sustainable LLMs on Code-Related Task

Shrinking AI code-helpers to save real energy — and checking if FLOPs even measures that.

LLMs used for coding tasks like spotting bugs or summarizing code are power-hungry, which is a problem for laptops and the planet. This paper studies 'knowledge distillation' — training a smaller, cheaper 'student' model to mimic a big 'teacher' model — but with an eye specifically on energy use rather than just raw computation. Using a tool called Morph, they balance accuracy against actual measured energy consumption. A key question they ask: does counting FLOPs (floating point operations) really predict real-world energy draw, or is it a misleading shortcut the field has been relying on?

Technical view

The study applies Morph, a many-objective optimization-based distillation framework, to compress LLMs for SE tasks (clone detection, vulnerability prediction, code summarization), treating measured energy consumption as an explicit optimization objective alongside accuracy. It empirically tests whether FLOPs correlates with actual energy draw, challenging FLOPs' status as the default efficiency metric in SE-LLM literature. Practitioners could replicate this by instrumenting hardware power measurement alongside distillation runs to validate FLOPs-based efficiency claims before deployment on constrained hardware.

arXiv · physics.comp-phBuildable

Agentic Porting, Construction and Initial Verification and Validation of Libraries within the Open Source Unified TRAnsient Multi-Phase Advanced Reactor simulation Kit (Outram Park) Part I: Thermal Hydraulics

AI coding agents are porting decades-old nuclear reactor simulation software into Rust, with humans checking the physics.

Nuclear reactor safety simulations rely on huge, old scientific codebases (like OpenFOAM) that are hard to maintain and modernize. Here, researchers used AI 'agentic' tools — AI systems that can autonomously rewrite code with a human checking their work — to translate parts of these simulation libraries into the safer, faster Rust programming language. The twist is that once AI handles the grunt work of translation, the hard part shifts to verification: making sure the new code still gives physically correct answers, tested against known benchmark problems like shock tubes. This is a real-world test of whether AI-assisted code porting is trustworthy enough for safety-critical engineering software.

Technical view

The authors use human-in-the-loop agentic code porting to translate OpenFOAM-derived libraries into Rust within the Outram Park reactor simulation kit, producing 'Outram-Foam' and building two-phase homogeneous-equilibrium (HEM) choked-flow solvers within the TAMPINES thermal-hydraulic library set (e.g., tampines-steam-tables). Preliminary verification/validation uses classic CFD benchmarks (lid-driven cavity, Sod shock tube) to confirm numerical fidelity post-port. The core claim is that V&V, not code generation, becomes the bottleneck when using AI agents for large-scale scientific code migration — a pattern replicable for other legacy Fortran/C++ codebases.

arXiv · cs.CVBuildable

REChart: Reasoning-Efficient Chart Editing with Large Reasoning Models

Teaching AI to edit charts from a picture and instruction — without letting it overthink and hallucinate.

Imagine showing an AI a picture of a bar chart and saying 'make the bars blue and add a title' — it needs to figure out the underlying code that drew the chart and modify it correctly. This requires careful visual reading, precise instruction-following, and working code generation. The researchers found that giving 'reasoning' AI models more thinking time doesn't always help — past a point, they start imagining details that aren't really there or get stuck in loops. Their fix, REChart, trains the model in two stages using 200,000 synthesized examples, giving feedback on intermediate reasoning steps so the model reasons just enough, accurately and efficiently, rather than endlessly.

Technical view

REChart addresses an empirically observed inverted-U relationship between chain-of-thought length and chart-editing accuracy in multimodal large reasoning models, where excessive CoT causes hallucinated visual grounding or redundant reasoning loops. The two-stage training framework applies process-level supervision over intermediate reasoning steps (not just outcome supervision) using a synthesized dataset of ~200k chart-instruction-code triples, targeting both editing fidelity and reasoning efficiency. Practitioners building chart-editing or vision-to-code agents could apply this process-supervision technique generally to control CoT length and quality in visually-grounded code-generation tasks.

arXiv · cs.SEBuildable

COMMITGUARD: Differential Slice Fuzzing for Commit-Induced Bug Detection

Fuzz-testing only the tiny slice of code a commit changed, using the old version as the truth check.

Every time a programmer commits a change, there's a risk of introducing subtle memory bugs that reviewers and tests miss. Running a full fuzzer — a tool that bombards code with random inputs to find crashes — on the entire program for every commit is far too slow. COMMITGUARD's trick is to compare the new version of a changed function against its old version, using the old behavior as a baseline for what's 'normal,' so it can focus testing precisely on the changed code and flag when new behavior suspiciously diverges. This turns bug-catching into an everyday part of commit review instead of a rare, expensive whole-program audit.

Technical view

COMMITGUARD performs commit-aware differential slice-based fuzzing: it extracts a program slice around the modified function, treats the pre-commit version's behavior as a baseline oracle, and fuzzes both versions to surface behavioral divergences indicative of memory-safety regressions (boundary, lifetime, initialization errors). This scopes both the search space and the correctness oracle to the diff itself, avoiding whole-program fuzzing cost and sidestepping hand-written oracles since divergence from prior behavior is the signal. Teams could integrate this into CI as a per-commit gate by building slice extraction plus dual-binary differential execution around an existing fuzzer like libFuzzer or AFL.

arXiv · cs.SEBuildable

SNIPTEST: Fuzzing Multi-Level Code Slices for Validating Vulnerabilities

Instead of fuzzing a whole codebase to check one warning, fuzz just a tiny compiled snippet around it.

Static analysis tools scan code and flag warnings like 'this might overflow a buffer,' but many turn out to be false alarms, and checking each by hand takes forever. Fuzzing can confirm real bugs, but running it against an entire large program for every warning is too slow. SNIPTEST instead cuts out just the small code slice around the warning, compiles that piece alone, and fuzzes it directly — expanding the slice if needed — to gather fast evidence about whether the warning is a real, exploitable problem, much quicker than testing the whole project.

Technical view

SNIPTEST is an execution-based warning-triage framework that generates compiled code slices centered on static-analysis warnings and fuzzes them directly rather than the whole binary, producing fast exploitability evidence. It progressively expands the slice's execution context when initial fuzzing is inconclusive, trading completeness for speed relative to whole-program directed fuzzing, which can take days for marginal coverage gains. Security teams could integrate SNIPTEST after tools like CodeQL or Clang Static Analyzer as an automated pre-filter before human triage review.

arXiv · cs.SEConceptual

Oracles That Cannot Fail: Anchoring and the Expectation That Moves With the Fault

Some software tests are secretly rigged to always pass — this paper hunts down that hidden flaw.

When testing software, you need an 'oracle' — a way of knowing what the correct answer should be, so you can tell if the program got it wrong. But sometimes both the expected answer and the actual result are computed from the same buggy internal state, so if something breaks, both numbers break together and still match — meaning the test can never catch the problem, no matter how many inputs you try. The researchers call this 'oracle anchoring' and define when an expectation is properly independent versus dangerously tied to the very code being tested. They tested this on a real deployed air traffic control simulator, deliberately injecting hundreds of bugs to see how many their method catches that were previously invisible.

Technical view

The paper formalizes 'oracle anchoring': a test oracle is unsound when its expected value is derived, directly or transitively, from the same mutated code under test, causing measurement and expectation to co-vary and mask faults regardless of input generation. They distinguish specification-anchored (externally fixed) from state-anchored expectations, enumerate additional channels for state-anchoring, and operationalize a predicate restricted to values flowing from the mutation target. Evaluated via mutation testing on a deployed ATC simulator across 4 modules, 12 property suites, and 366 mutants, the technique surfaces additional mutants previously undetected due to anchored oracles — giving practitioners a concrete check for auditing property-based test suites before trusting mutation-testing results.

arXiv · cs.SERunnable

Graphectory Viewer: A Tool for Process-Centric Analysis of Agentic Software Trajectories

A visual X-ray for watching AI coding agents think, step by step, across thousands of runs.

When an AI agent works through a coding task, it produces a messy log of thoughts, actions, and observations that's hard for a human to make sense of. Graphectory Viewer turns these raw logs into interactive graphs showing the phases of problem-solving — like exploring, editing, testing — so researchers can click into any step to see exactly what the agent was doing. It also shows flow diagrams comparing what successful runs looked like versus failed ones, across many agent frameworks at once. As AI agents get more autonomous, understanding how they succeed or fail — not just whether they did — becomes essential for improving them.

Technical view

Graphectory Viewer is a web-based visualization tool that converts heterogeneous agent trajectory logs from multiple agent frameworks into the 'Graphectory' representation — phase-aware graphs linking low-level execution (thoughts/actions/observations) to higher-level behavioral structure. It supports interactive graph construction, node-level inspection, search/filtering across large trajectory corpora, and Sankey-style summaries of phase transitions to compare successful vs. failed runs. Researchers evaluating coding agents on benchmarks like SWE-bench could use it to diagnose failure modes and behavioral drift at scale rather than relying solely on pass/fail metrics.

arXiv · cs.SEBuildable

Grounding AI Agents in Contracts: An Empirical Evaluation of Spec-Driven Test Generation

Making AI write down a program's rules before writing tests for it — and it catches more real bugs.

When you ask an AI to write software tests directly, it often misses tricky edge cases because it doesn't deeply understand what the code is supposed to guarantee. This paper's fix is to make the AI first explicitly write out the code's 'contract' — its preconditions (what must be true going in), postconditions (what must be true coming out), and undefined behaviors — before generating any tests. It's like making a detective write down their theory of the crime before questioning suspects, so their questions are sharper. Tested on real production bugs from Google, this 'spec first' approach found measurably more actual bugs than just asking the AI to write tests directly.

Technical view

Spec-Driven Test Generation inserts an intermediate step where an LLM agent is prompted to reason about and explicitly document a function's pre-conditions, post-conditions, and undefined behaviors as a semi-formal specification, which then scaffolds subsequent test generation. Evaluated on a corpus of production bugs from Google, the spec-driven agent achieved a statistically significant 9.8 percentage point improvement in bug detection rate (p=0.0352) over direct test generation. This is replicable via simple prompt-chaining — a spec-extraction step followed by a test-generation step — without any model retraining, making it directly adoptable in existing LLM test-generation pipelines.

arXiv · cs.SEConceptual

A Multi-Surface Consistency Audit of Software Citation Metadata

Software projects describe themselves differently in five places at once — and it's a mess.

Research software (the code scientists use to run experiments) gets 'described' in many separate spots: a citation file in the code repository, an official archive record, a DOI registry entry, a package listing, and the README. The problem is nobody had checked whether these self-descriptions actually agree with each other — different tools and citation systems often read only one of these spots, so if they disagree, credit for the work can get scattered or lost. The researchers audited 117 real open-source science software projects, comparing each one's metadata across all these surfaces to see how often they matched. This matters because proper credit and traceability for the software behind scientific results depends on this metadata being consistent.

Technical view

The authors conduct an empirical consistency audit across five metadata 'surfaces' (CITATION.cff, archival deposits like Zenodo, DOI registry records, package registry metadata, and README text) for 117 research software projects (87 HPC/quantum computing repos + a 30-project registered corpus), measuring agreement rates between surfaces on fields like title, authors, and version. The core contribution is a direct, previously unmeasured quantification of metadata fragmentation risk affecting citation tracking and provenance systems that sample only a subset of surfaces. Practitioners building citation-graph tools, software heritage indexes, or automated credit-attribution agents can use these findings to identify which surface pairs are most failure-prone and prioritize reconciliation or single-source-of-truth tooling accordingly.

arXiv · cs.LGBuildable

From Abductive Explanations to Global Logical Rules for Node Classification in SGCs

Teaching an AI to explain its graph predictions with clean, general rules instead of messy one-off reasons.

Graph neural networks are AI systems that make predictions about nodes in a network (like classifying people in a social graph), but they're famously hard to explain. Earlier explanation methods pull out little example subgraphs that justify each prediction, then try to turn those into general rules — but those examples are often cluttered with details specific to just one node, making the resulting rules sloppy. This paper instead computes, for each node, the smallest possible set of clues (feature values) that's truly necessary to keep the AI's prediction the same, stripping away the noise before generalizing. Those minimal clues are then fed into decision trees, a simple and readable AI model, to extract clean, global logical rules that describe how the network actually reasons. This matters because trustworthy AI in high-stakes settings requires explanations that are both accurate and genuinely understandable, not just locally plausible.

Technical view

The method targets Simple Graph Convolution (SGC) networks and computes minimal abductive explanations per node — the smallest sufficient subset of node-feature pairs that preserves the predicted class — as a cleaner intermediate representation than the redundant explanatory subgraphs used by prior work like LogicXGNN. These minimal explanations are used to train decision trees, from which global logical rules are extracted, aiming to reduce node-specific redundancy and improve rule generality. This offers a reproducible pipeline for practitioners: compute abductive explanations via existing SAT/ILP-based minimality solvers, then train interpretable surrogate models to derive auditable, human-readable classification rules for GNN behavior on SGC architectures.

arXiv · cs.LGBuildable

Backward through Time, Algebraically

Rebuilding the rulebook for teaching AI to follow time-based instructions using math, not shortcuts.

Linear temporal logic is a formal language for stating rules like 'always do X before Y' — traditionally used with strict true/false answers. But modern AI systems (like neural network controllers) don't work in strict true/false; they deal in degrees and probabilities, so researchers need a version of these time-rules that can be 'graded' and used as a training signal via calculus-style differentiation. The tricky part is there are many competing ways to make these rules gradable, and existing tools lock you into just one choice upfront. This paper builds a flexible engine that lets you plug in whatever grading scheme you want and still get differentiable, trainable feedback — instead of committing to one fixed system. This matters for anyone training AI agents or controllers to reliably satisfy time-based goals (like 'never crash, then eventually reach the target').

Technical view

The paper addresses the problem of differentiable semantics for Linear Temporal Logic (LTL) when applied to softly-valued systems (neural policies, adaptive controllers), where formula satisfaction becomes a differentiable training signal rather than a boolean. Instead of a shallow embedding committed to one semantic algebra, the authors build an algebra-generic evaluation engine (framed via functional programming abstractions) that supports pluggable differentiable semantics and remains amenable to automatic differentiation. This is directly useful for researchers building neurosymbolic or temporal-logic-guided RL/control systems who want to experiment with or swap semantic algebras (e.g., product t-norms, Gödel fuzzy logic) without reimplementing the evaluation engine.

arXiv · cs.LGBuildable

Certified but Private: Scalable Zero-Knowledge Proofs for Neural Network Guarantees

Proving your AI model is fair and robust — without ever showing anyone the model.

Companies often can't share their AI model's internal parameters because they're valuable trade secrets, yet regulators or users may need proof that the model behaves safely — like being 'robust' (not easily fooled) or 'fair.' This paper builds a system called PANDA that uses zero-knowledge proofs, a cryptographic trick that lets you prove a statement is true without revealing the secret information behind it, to certify these properties about a neural network without exposing its weights. It works by adapting an existing robustness-checking method (CROWN) and inventing a lightweight way to cryptographically prove the math behind checking non-linear parts of the network (like activation functions), keeping the proofs fast and practical. This matters because it could let auditors verify AI safety and fairness claims in legally sensitive situations while companies keep their models confidential.

Technical view

PANDA combines the CROWN linear-relaxation robustness certification framework with zero-knowledge proofs, contributing a novel algorithm for proving linear relaxation bounds over non-linear activation layers that yields lightweight, practical ZK proofs rather than prohibitively expensive general-purpose circuit proofs. The system generates proofs of local robustness (and reportedly fairness properties) for neural networks while keeping model parameters private from the verifier. This is significant because prior ZKP approaches to neural network verification scale poorly; PANDA's contribution is specifically the efficient proof construction for the non-linear bound-propagation step, making it a candidate building block for auditors or regulators needing verifiable claims about proprietary models in safety-critical or legal-compliance contexts.

arXiv · physics.comp-phRunnable

Validating direct solvers for Newton's gravitational N-body problem, and the systematic comparison between IEEE floating point and Posits

Testing whether a quirky new number format beats standard computer math for simulating orbiting planets.

Simulating gravity between many bodies (like planets or stars) over time is a classic 'chaotic' problem where tiny rounding errors in a computer's number system can snowball into wildly wrong predictions. Computers usually represent numbers using the IEEE floating-point standard (the fp16/fp32/fp64 you might have heard of), but there's an alternative format called Posits that some claim represents numbers more efficiently. This study runs the same gravity simulations using several different number formats — different precisions of standard floating point and two versions of Posits — and compares them against near-perfect 'arbitrary precision' math to see which is most accurate and how fast each runs. The finding: the lowest-precision formats are simply too imprecise for this kind of physics, while higher-precision options are needed depending on whether you want single close-encounter accuracy or just statistical trends. This matters for anyone building large-scale astrophysical simulations who has to trade off computing speed against numerical accuracy.

Technical view

The authors benchmark direct N-body gravitational solvers across IEEE-754 formats (fp16, bfloat16, fp32, fp64, fp128) and two Posit (type III unum) implementations, validating each against arbitrary-precision arithmetic as ground truth, using hardware/compiler support for fp64 and software implementations elsewhere. Key finding: half-precision formats (fp16, bfp16, Posit<16,1>) are unusable for Newtonian N-body integration due to insufficient precision, fp32/Posit<32,2> are viable only for statistical ensemble-level results (not individual close encounters), and 64-bit formats are needed for high-fidelity trajectories. This gives computational astrophysicists concrete precision/performance tradeoff data for choosing number formats in large-scale N-body codes, and is among the first systematic head-to-head comparisons of Posits against IEEE floats for this chaotic dynamical system.

arXiv · cs.SEBuildable

LadderTeam: Dual-Agent Laddering Elicitation Framework

Two AI agents interview each other's users to dig up what people actually want from an app.

When building software, developers need detailed feedback from users about what they really need — a technique called 'laddering' interviews does this well by repeatedly asking 'why' to drill down to root motivations, but it's slow, expensive, and requires skilled human interviewers. LadderTeam automates this using two AI agents talking to each other: one acts as an active interviewer probing a user (or a wireframe/mockup) with different questioning strategies, while presumably another plays a supporting or evaluating role. This turns a manual, hard-to-scale process into something reproducible and automatic. This matters because it could let small teams get rich, structured user feedback on their designs without hiring professional interviewers for every study.

Technical view

LadderTeam automates laddering-technique UX interviews (which elicit means-end chains from concrete product attributes to abstract user values) via a dual-agent LLM architecture, where an active interviewer agent applies one of three probing strategies against wireframes to extract detailed, actionable requirements. The framework is presented as open and reproducible, targeting the traditional laddering interview's scalability bottleneck (manual, time-and-cost-intensive, dependent on interviewer skill). Practitioners in requirements engineering or UX research could adopt this to run automated, repeatable elicitation studies on wireframes at scale, using the paper's probing strategies as a starting design space for prompt engineering their own interviewer agents.

arXiv · cs.SEBuildable

ORCA: Observability-Grounded Program Repair for Microservice Incidents

An AI that reads a service outage's system logs and writes the code fix itself.

When a microservice-based app (software split into many small interacting services) breaks, engineers usually diagnose the problem using operational data like logs and metrics, but the AI tools that try to auto-fix bugs usually start from a bug report or a failing test instead — missing the messy telemetry data that actually reveals what went wrong. ORCA closes this gap: it compares telemetry from a failure against telemetry from when things worked correctly to build a 'fault signature' describing what's different, uses that to guess which code or configuration is likely broken, then has AI agents generate candidate code patches. Crucially, it checks each patch not just by running tests but by replaying the actual telemetry to see if the fix truly resolves the observed failure pattern. This matters because it could make automated incident response for complex cloud systems far more grounded in real operational evidence rather than guesswork.

Technical view

ORCA is an observability-grounded automated program repair (APR) pipeline for microservice incidents: it distills paired failure/reference telemetry into a 'fault signature,' uses it to localize candidate code and deployment-config sites, generates unified-diff patches via repair-graph and exploration agents, and validates candidates with a Telemetry-Grounded Patch Verifier that separately checks patch validity, syntactic/semantic correctness, test-oracle integrity, and telemetry replay fidelity. Evaluated on a 575-case benchmark, ORCA reportedly outperforms existing APR baselines that rely only on issue reports or failing tests. This is a concrete architecture (signature extraction → localization → multi-agent patch generation → multi-stage telemetry-based verification) that practitioners building AIOps/SRE automation could adapt, particularly the telemetry-replay verification step as a novel patch-validity signal beyond unit tests.

arXiv · cond-mat.mtrl-sciRunnable

PowderLine: a programmatic powder diffraction analysis application

Turning fiddly crystal-analysis scripts into a single reusable recipe file for robot-run labs.

Scientists studying powdered crystalline materials use X-ray diffraction data and a technique called Rietveld refinement to extract detailed structural information, but doing this well requires expert know-how and usually means writing custom one-off scripts for each experiment. As labs increasingly run automated, self-driving experiments that generate huge amounts of diffraction data, they need this analysis to happen automatically and produce clean, structured output rather than requiring a human expert each time. PowderLine solves this by letting users define an entire refinement process as one declarative 'recipe' file — following a standardized, checkable template — which the software then validates and runs itself, spitting out structured results ready for further automated use. This matters because it helps 'self-driving labs' (automated research facilities) run consistent, reliable materials analysis at scale without a human in the loop for every sample.

Technical view

PowderLine is a Python application that encapsulates whole-pattern (Rietveld) refinement workflows into a single declarative, schema-validated recipe format, which it executes via underlying refinement software and returns as structured, machine-readable results. The recipe is described as an all-inclusive, versioned, machine-readable/-writable specification, enabling programmatic, reproducible refinement runs suited to high-throughput and autonomous/self-driving-lab pipelines. Practitioners can adopt PowderLine to replace bespoke refinement scripts with a standardized schema-driven interface, easing integration into automated experimental pipelines and enabling version-controlled, auditable refinement configurations across large sample sets.

arXiv · cs.LOConceptual

Simplicial Actions for Distributed Protocols

Teaching computers to reason about who-believes-what when the group is out of sync.

Simplicial complexes are a way of drawing a group of processes or agents as connected shapes (points, edges, triangles) where each shape encodes what a subset of them can jointly know or believe. This paper extends that toolkit with 'action models' — formal descriptions of events, like a message arriving, that change what everyone believes, including cases where new information forces someone to revise a belief they held before. The real-world problem is distributed computing: separate processes coordinating with incomplete information need rigorous tools for reasoning about knowledge and belief change, not just intuition. The authors explicitly connect this machinery to distributed task computability, the theory of what coordination problems are solvable at all. It matters because it puts belief-revision reasoning about distributed systems on firmer mathematical footing.

Technical view

The paper builds on simplicial semantics for modal logic, where facets of a simplicial complex encode local epistemic states of distributed processes, and extends action models — previously developed mainly for knowledge — to simplicial belief models that support revision. It integrates prior results on belief in simplicial complexes and belief-revision semantics, generalizing action models beyond earlier task-computability-focused treatments in the literature. The framework explicitly links to distributed task computability, offering a way to model how beliefs among asynchronous processes update after communication events. Researchers formalizing fault-tolerant distributed protocols via epistemic/doxastic logic, rather than purely operational semantics, can build directly on this action-model construction.

arXiv · cs.SERunnable

ModBench: A Pipeline for Building Modelica Benchmark Datasets Mined from Library Repositories

A robot archive-digger that turns years of engineering-model commits into a ready-to-use dataset.

Modelica is a language engineers use to build simulations of physical systems, like cars or power plants, and researchers studying how these models evolve have lacked a clean dataset to work with — just messy raw commit histories. ModBench fixes this by automatically sifting through a repository's Git history, keeping only meaningful human-made changes, pulling out the parts of each model that can actually be simulated, and converting them into a standard, comparable format. Applied to the official Modelica Standard Library, it produced over 85,000 model snapshots spanning the library's entire history, each traceable back to its original source and commit. This gives researchers a solid, reusable foundation for studying how complex engineering models are built, debugged, and improved over time.

Technical view

ModBench is a mining pipeline for Modelica libraries that filters Git commit history to human-authored, relevant revisions, extracts simulation-eligible classes, and normalizes them into canonical representations for cross-version comparability. Applied to the Modelica Standard Library, it produced 85,562 distinct class snapshots spanning the full commit history since language v3, each traceable to original models and Git metadata. The released dataset, API, and generation pipeline let researchers benchmark model-evolution analyses, regression detection, or LLM-based Modelica code generation against a standardized, versioned corpus instead of ad hoc scraped repos. Practitioners can extend the pipeline to other Modelica libraries to build comparable domain-specific benchmarks.

arXiv · cs.SEBuildable

The Working Set of a Coding Agent: Coherence Debt in Repository-Scale Tasks

Coding AIs fail exactly where a project's memory and prompt both go blank — and nowhere else.

When an AI assistant edits code across a large codebase, it needs many connected facts to stay consistent — like knowing a test expects a certain function name or that a config file must match an import elsewhere. This paper studies what happens when some of those facts aren't available, either because they're missing from what's shown to the AI or because it never learned them during training. The researchers systematically hid and revealed different facts and injected errors across seven AI models and five coding-agent setups. They found that when a fact is missing from both sources the AI reliably fails at that exact point, regardless of which model is used, and that showing a fact works just as well whether it's near or far from the edit. It matters because it suggests reliability isn't about smarter models — it's a solvable engineering problem of making sure the right facts are visible when needed.

Technical view

The paper formalizes repository-scale editing as reconstructing a coupled-fact graph, where each required fact for an edit must come from recent context or parametric memory; facts absent from both constitute 'coherence debt.' Across seven LLMs and five agent harnesses, with faults injected and context/memory channels independently withheld, they show total failure when both channels are empty, and that supplying a missing fact restores success equally well regardless of its distance from the edit site — availability, not locality, predicts outcomes. When a library rename invalidates memorized API knowledge, all seven models fail identically, missing the same tests, indicating a shared failure mode rather than model-specific weakness. This gives practitioners a diagnostic framework — measure which facts are covered by context versus memory — for deciding what to retrieve or inject rather than relying on bigger context windows alone.

arXiv · cs.SEConceptual

The Specification Paradox: Rethinking Requirements Engineering in the Age of AI

AI didn't kill the hard part of software — it moved the pain to figuring out what to build.

There's a common hope that AI coding tools will let anyone skip the hard work of software development just by describing what they want. This paper argues that's wishful thinking in new clothes: AI reduces the effort of typing code, but the real difficulty of software was never the typing — it was understanding the problem and keeping the definition of 'correct' consistent as things change. The authors propose shifting focus to Specification-Driven Development, where writing precise, validated specifications becomes the central skill, since AI can turn a good specification into code far more reliably than a vague one. It matters because it reframes what teams should invest in: not just faster code generation, but better ways to capture, refine, and evolve requirements.

Technical view

This is a position paper arguing that LLM-driven code generation shifts the software engineering bottleneck from implementation to requirements elicitation, specification authoring, and validation — a move the authors term Specification-Driven Development. It contends that AI's productivity gains are conditional on upstream specification quality, making Requirements Engineering newly central rather than obsolete, and discusses implications for how RE practices, artifacts, and roles should adapt. No empirical study is described in the abstract; it reads as a conceptual contribution meant to reorient research and practice agendas around specification quality as the binding constraint. Practitioners can use this as a framing lens for deciding where to invest tooling effort, e.g., specification validation and requirements traceability, rather than solely in code-generation quality.

arXiv · cs.SEConceptual

Factors Impacting Developer Efficiency: Results from an Adaptive Longitudinal Study

A year-long check-in with 27 developers reveals what actually slows them down, and it isn't the code.

This study asks a deceptively simple question: what really gets in the way of developers doing their jobs well, and does that change over time? Instead of a one-off survey, researchers followed 27 developers in a consulting environment for over a year, checking in with surveys twelve separate times and running in-depth interviews to hear things in their own words. They found that the biggest, most persistent obstacles weren't about coding skill or tools at all — they were organizational, like waiting on other teams for approval — and these problems stayed remarkably stable throughout, with technical issues playing a secondary role. It matters because it pushes back on the assumption that developer productivity is mainly a technical problem solvable with better tools; often it's a people-and-process problem instead.

Technical view

The study applies the Adaptive Developer Efficiency Monitoring Method (ADEMM) as a mixed-methods longitudinal design, combining twelve waves of periodic surveys with eighteen semi-structured interviews across 27 external developers in a consulting/professional-development context, analyzed via statistical and thematic methods. The dominant, structurally stable bottlenecks identified were organizational dependencies and waiting for external validation, with technical factors appearing as secondary and more variable contributors. This provides an empirical basis for prioritizing process and organizational interventions, such as reducing approval latency or cross-team dependency bottlenecks, over purely technical tooling investments when optimizing developer efficiency. Researchers could replicate the ADEMM protocol in other consulting or outsourced-development contexts to test how well the findings generalize.

arXiv · cs.SEBuildable

ADEMM: A Longitudinal Method for Monitoring Developer Efficiency in Industry

A repeatable playbook for tracking why developers you don't directly manage are struggling — over time.

Companies often want to understand what's slowing down developers, but most measurement approaches are a single snapshot — one survey, one point in time — which misses how problems shift and evolve. This is especially tricky when the developers being studied work for another company, like consultants or trainees, rather than being direct employees. This paper introduces and tests ADEMM, a structured method combining regular surveys and periodic interviews to continuously track developer-efficiency barriers, refined over five rounds of trial and improvement together with the organization using it. It matters because it gives organizations a reusable, evidence-based recipe for ongoing monitoring rather than one-off diagnostics, which is especially useful when managing external or partner teams.

Technical view

ADEMM is a longitudinal monitoring method designed via Design Science Research and Action Design Research, iteratively refined across five cycles through a mixed-method study with 27 external developers over twelve survey waves plus eighteen semi-structured interviews, co-evaluated with the host organization. It targets contexts where the monitoring organization lacks direct employment authority over the developers, such as consulting or professional-education settings, combining recurring quantitative surveys with qualitative interviews to detect how efficiency barriers emerge and shift over time. This provides a transferable protocol — survey instrument, interview cadence, iterative refinement process — that organizations or researchers can adopt to build their own continuous developer-experience monitoring programs. The companion paper reports the empirical results this method produced.

arXiv · cs.SEBuildable

Operationalizing the EU AI Act in Agile Software Development: A Guideline-Based Approach

Turning EU AI Act legalese into checklist items your Scrum team can actually tick off.

The EU AI Act is a new law requiring companies that build or deploy AI systems to document their work, manage risks, and keep humans in the loop — but it's written in dense legal language that doesn't map onto how agile teams actually work, in short sprints with things like a 'Definition of Done' or sprint reviews. This paper builds a practical guideline that translates the law's abstract requirements into concrete actions agile teams can slot into their existing workflow, and explains the method used to do that translation. They rated each article of the law by how directly relevant and urgent it is, using a red/yellow/green priority scheme, then mapped the important ones onto specific agile practices. It matters because it gives real teams a way to stay legally compliant without abandoning agile development, and the same translation method could be reused for other regulations.

Technical view

Following Design Science Research methodology, the authors assess each article of the EU AI Act along three dimensions and classify them with a traffic-light scheme, then map high-priority articles onto concrete artifacts and ceremonies within agile practice, such as Definition of Done, Sprint Review, and working agreements. The output is an evaluated, actionable compliance guideline plus a documented translation method intended to be reusable for mapping other regulations onto agile workflows. This is directly applicable for teams shipping AI features under EU jurisdiction who need concrete documentation, risk-management, and human-oversight artifacts rather than legal-text interpretation; the traffic-light prioritization method itself is a reusable technique for future regulatory-to-practice translation work.

arXiv · cs.GTConceptual

Solving Streett and Emerson-Lei Games with Universal Trees

Cracking open the math trick behind the fastest known solutions to a whole family of infinite-game puzzles.

Imagine two players taking turns forever on a game board, and you need pure logic to determine who has a winning strategy — these 'infinite games' show up in computer science to model and verify whether a program will always eventually do what it's supposed to. A breakthrough a decade ago showed one type of these games, called parity games, could be solved surprisingly fast using a structure called a 'universal tree.' This paper extends that trick to two harder, more general types of games, Streett and Emerson-Lei games, which previously could only be solved fast by awkwardly converting them into the simpler type first. They work out directly how universal trees interact with these games' underlying structure, producing a genuinely faster algorithm rather than a roundabout conversion. It matters because these games underlie automated verification tools that check whether critical software behaves correctly forever, so faster solving means checking bigger, more complex systems.

Technical view

The paper extends the universal-tree framework — previously understood mainly for parity and Rabin games — to give a direct, non-reduction-based algorithm for solving Streett and Emerson-Lei games, refuting the assumption that universal trees only apply to games admitting memoryless winning strategies. It characterizes how universal trees interact with Zielonka trees, the structures encoding acceptance conditions for these more general omega-regular game classes. The resulting algorithm solves Streett games with n vertices, m edges, and k pairs in O(mk·log(k)·k!·|U(n,k)|) time, avoiding the overhead of reducing to parity games first. Researchers in formal verification and model checking can build on this to implement faster direct solvers for Streett/Emerson-Lei acceptance conditions used in reactive synthesis and temporal-logic verification tools.

arXiv · cs.SEBuildable

DCI: Dependency Confidence Index for Assessing Open-Source Dependency Trustworthiness

A trust score that tells you whether that open-source library you're about to add is actually safe.

Every modern app pulls in dozens of outside code libraries, and picking a shady or poorly-maintained one is a real security risk — this is the software supply-chain problem. The researchers built the Dependency Confidence Index, a single score combining nine trust signals like security history, code quality, and how healthy the project's community is. They figured out how much each signal should count by surveying developers and reviewing prior research, then automated the actual measuring with tools like SonarQube and GitHub data. The goal is a quick, evidence-based number developers can check before adopting a package instead of guessing.

Technical view

DCI is a formative composite index built from nine trust factors weighted via an Analytic Hierarchy Process survey of ten developers plus a systematic literature review, implemented as 12 automated GQM-derived measurements using SonarQube, GitHub APIs, and OpenSSF Scorecard, run through a containerized evaluation pipeline. A pilot on 92 popular PyPI packages showed only moderate agreement with OpenSSF Scorecard, suggesting DCI captures somewhat different trust signals rather than duplicating existing tools. Practitioners could replicate the pipeline against their own dependency sets or reweight factors for domain-specific risk tolerance.

arXiv · cs.SEConceptual

Towards Risk-free AI Agent Deployment

To trust an AI agent in production, watch every step it took, not just what it said.

AI agents that reason, call tools, and act on their own are increasingly running real business processes, but a wrong step can cause security, compliance, or functional failures. The authors argue the key to catching these failures is the 'trajectory' — the full recorded trail of what the agent thought, did, and observed along the way, since many bugs are invisible unless you look at that trail. They lay out why testing agents is uniquely hard (agents behave differently each run, and there's no simple way to check if an output is 'correct'), and argue for treating agent debugging — tracing exactly which step caused a failure and fixing it — as its own research field. It matters because without this, companies are deploying unpredictable software with no reliable way to catch its mistakes.

Technical view

This is a position paper framing agent trajectories — the sequence of reasoning steps, tool invocations, and environment observations — as the fundamental unit for both testing and debugging LLM-based agents. It catalogs testing challenges specific to agents: the oracle problem (no ground truth for 'correct' behavior), non-determinism across runs, trajectory validation, and lack of adequacy metrics analogous to code coverage. On the debugging side it points toward automated failure attribution (pinpointing which trajectory step caused a failure) and automated repair/self-evolution as concrete research directions worth building tooling around.

arXiv · cs.LOBuildable

SATisfying the High School Identities but not Wilkie's Identity

A computer brute-force search just closed a decades-old puzzle about the rules of algebra.

Tarski asked whether the everyday rules of algebra you learn in high school (like distributing multiplication over addition) are powerful enough to prove every true statement about numbers — turns out they aren't, because mathematician Wilkie found one true identity those rules can't derive. This paper asks a follow-up question: is there some small abstract number-like system that obeys all the high-school rules but still breaks Wilkie's special identity? The authors turned this into a giant logic puzzle and fed it to a SAT solver (software that finds solutions to yes/no constraint problems), exhaustively checking every possible small system up to size 11 and confirming none exists, with an independent double-check. Along the way they also stumbled on a brand-new 12-element example that breaks the rules in a genuinely different way from the one mathematicians already knew.

Technical view

The paper resolves an open question tied to Tarski's High School Algebra Problem by encoding the search for an 11-element algebra satisfying the High School Identities while refuting Wilkie's identity as a SAT instance, showing via exhaustive solver search (with independent verification) that no such algebra exists. As a byproduct they found a new 12-element countermodel non-isomorphic to the previously known witness, expanding the known landscape of such algebras. Others could replicate or extend this by adapting the SAT encoding to search larger algebra sizes or different identity sets.

arXiv · cs.PLBuildable

Mechanizing Choreographic Programs and Hoare Logic with State Transformers

Formally proving distributed programs can never deadlock, without getting bogged down in variable bookkeeping.

Choreographic programming lets you write one program that describes an entire multi-computer conversation at once, and a compiler automatically splits it into the separate pieces each computer runs — with the neat guarantee that the pieces can never get stuck waiting on each other forever (deadlock). This paper is about mechanizing that idea, meaning formally verifying it inside a proof assistant (software that checks mathematical proofs line by line) along with Hoare logic, a standard way of reasoning about what a program's code actually does. The fiddly part of these proofs is usually tracking how variable names get substituted around, so the authors borrow a recent technique to sidestep that bookkeeping entirely and keep the proof focused on what's actually distributed-specific. This matters because it makes trustworthy formal guarantees about distributed systems easier to build and check.

Technical view

The work mechanizes choreographic programming together with a Hoare logic for reasoning about choreography programs, using a state-transformer-based technique recently proposed by Thiemann to model deadlock-freedom without the usual overhead of variable binding and substitution machinery. By separating distributed-specific operations from standard local operations already handled in non-distributed program semantics, the mechanization stays concise while preserving deadlock-freedom-by-construction guarantees. Researchers building formally verified distributed systems tooling could adopt this state-transformer approach to keep their own choreography mechanizations lean.

arXiv · cs.SERunnable

Comparing the Quality of Code Generated by Vibe Coding Tools

Three AI app-builders were told to build the same thing — their code quality turned out very different.

'Vibe coding' tools let you type a single sentence and get a working web app back, and this study checks how good that generated code really is under the hood, not just whether it runs. The researchers gave the same prompt to three popular tools — Lovable, v0, and Replit — three times each, producing nine apps total, then ran them through SonarQube, a standard code-quality scanner that flags bugs, messy duplicated code, and overly complex logic. Early results show each tool has its own quality 'fingerprint': for instance Lovable tends to rack up many minor code smells rather than serious bugs. This matters because as more real software gets built this way, knowing the maintainability trade-offs of each tool helps developers choose wisely and know what to clean up.

Technical view

The study benchmarks structural code quality across Lovable, v0, and Replit by generating three independent projects per tool from a single prompt (nine web apps total) and running static analysis via SonarQube, capturing issue counts, severity distribution, estimated remediation effort, cyclomatic/cognitive complexity, and duplication. Preliminary findings show distinct tool profiles — e.g., Lovable skews toward high-density, low-severity code smells rather than fewer, higher-severity issues. Practitioners evaluating vibe-coding tools for production use could reuse this SonarQube-based methodology to audit maintainability before adopting a tool at scale.

arXiv · cs.SEConceptual

Implicit, Yet Impactful: Understanding Hidden Dependencies in Java Projects

Java projects quietly depend on libraries they never actually declared using.

When you build a Java project, package managers automatically pull in a web of dependencies based on what you explicitly ask for. But sometimes your code directly uses a library that got pulled in indirectly and was never officially declared as a dependency — an 'implicit dependency' hiding in plain sight. That's risky because you don't control its version, so it could change or disappear without warning, creating hidden security and maintenance problems. This is the first study to actually measure how common and how impactful these silently-used-but-undeclared dependencies are across real Java projects.

Technical view

The paper introduces implicit dependencies as a distinct category from direct and (unused) transitive dependencies — libraries actively referenced in project code but never explicitly declared, leaving their resolved versions outside developer control. It presents the first quantitative empirical study measuring prevalence and impact of this phenomenon in Java projects, likely via static analysis of import/usage versus declared manifest entries. Build-tool maintainers could use these findings to motivate stricter dependency-declaration linting (akin to Maven's dependency:analyze) to surface implicit dependencies automatically.

arXiv · cs.SERunnable

Validating HTTP Semantics in REST APIs With Constructed Call Sequence Scenarios

A smarter fuzzer builds sneaky sequences of web-API calls to catch APIs that break HTTP's own rulebook.

REST APIs — the way most web services talk to each other — are supposed to follow HTTP's rules about things like status codes and caching, and breaking those rules makes APIs confusing or buggy in ways that can cause real damage. The researchers extended an existing automated API-testing tool called EvoMaster with nine new checks specifically for HTTP rule violations. After the tool's normal random testing phase finishes, a second phase takes those generated tests and deliberately rearranges them into new call sequences designed to expose exactly these HTTP violations. When tested on nine sample APIs with bugs intentionally planted in them, the technique caught every single one.

Technical view

The authors extend the EvoMaster fuzzer with nine new oracles targeting HTTP semantics-level faults, then add a post-fuzzing phase that reuses the N generated test cases as seeds to synthesize new call sequences specifically crafted to probe each oracle's HTTP property (e.g., idempotency, caching, status-code correctness). Evaluated on 9 artificial REST APIs with injected faults, the approach achieved full detection of all injected faults. Since it builds on the open-source EvoMaster fuzzer, practitioners can plug their own REST API into the extended tool to run this HTTP-semantics validation directly.

arXiv · cs.SEConceptual

Software Engineering for AI-driven Building Operation

AI controlling your building's heating can't be tested like normal software — mistakes there are physical and permanent.

AI systems that automatically control heating, cooling, and other building operations promise big energy savings through smarter, predictive decisions. But current best practices for building and testing AI software assume that if something goes wrong, the worst case is a bad user experience — in a real building, a bad AI decision can waste energy that's gone forever, make occupants uncomfortable, or physically wear down equipment faster, none of which you can just 'undo.' Drawing on two research projects that combined civil engineering with computer science, the authors argue that existing software engineering practices for AI need to change to account for these physical, lasting, irreversible consequences. This matters as AI increasingly gets embedded in the physical infrastructure around us, not just apps on a screen.

Technical view

This is a position/experience paper arguing that standard Software-Engineering-for-AI (SE4AI) practices, built around digital-environment failure models, don't account for cyber-physical deployment contexts like building automation, where errors cause irreversible energy waste, comfort violations, or equipment degradation even though building systems are typically fault-tolerant and rarely safety-critical. Grounded in two interdisciplinary civil-engineering/computer-science research projects on AI-driven building operation optimization, the authors identify specific missing perspectives in current SE4AI methodology. It's aimed at motivating new testing, validation, and deployment practices tailored to physical, irreversible-consequence AI control systems rather than proposing a concrete technique.

BIO

Biology

95 new
arXiv · cs.CVBuildable★ flagship

Unsupervised Learning of Cell Instances with Generative Routing Pyramids

Find and describe every cell in a microscope image without anyone labeling a single one.

Analyzing microscope images usually means someone hand-labels thousands of cells to train a detector, then runs separate steps to segment each cell and describe its type — tedious and annotation-hungry. This method skips the labels entirely: it learns to reconstruct each image by routing pixels to a small set of sparse 'source' points through a coarse-to-fine pyramid, and in doing so it naturally discovers which pixels belong to which cell. Those pixel-to-source groupings become the instance masks, while each source's compressed code captures the cell's shape and morphology, giving you both segmentation and a phenotype fingerprint at once. Because it's unsupervised, it works across different cell shapes and imaging setups without retraining on new annotations. It matters because manual labeling is the main bottleneck in scaling up biological image analysis.

Technical view

The method performs unsupervised cell instance segmentation and phenotypic representation by reconstructing each image with a coarse-to-fine 'generative routing pyramid' that associates pixels with spatially sparse latent sources. The pixel-to-latent assignments directly yield instance masks, while the per-source latents encode morphology usable for phenotypic classification — unifying segmentation and representation in one generative pass rather than the usual detect-then-featurize pipeline. Reported results show competitive instance-segmentation performance across diverse cell morphologies and imaging modalities plus phenotype/gene-related analysis, all without manual annotations. Practitioners could apply it to unlabeled microscopy corpora to bootstrap masks and morphology embeddings, or adapt the sparse-routing reconstruction objective as an annotation-free pretraining stage.

arXiv · q-bio.BMConceptual★ flagship

Recovering protein conformations from single-particle cryo-EM data via indirect shape matching gradient flows

Read a protein's shape straight from blurry microscope snapshots—skipping the usual 3D map.

Cryo-electron microscopy freezes millions of copies of a protein and photographs them from random angles, but each photo is a noisy, flattened shadow rather than a clear 3D picture. Normally scientists first stitch these shadows into a fuzzy 3D density map and then guess where the atoms sit; this work skips that middle step and fits the protein's atomic skeleton to the raw photos directly. They start with a rough template of the backbone and gently bend and twist it—like reshaping a wire model—until the shadows it would cast match the real photos. Because proteins wiggle between different shapes to do their jobs, being able to catch those shape changes matters for understanding how they work and how drugs might target them. Doing it in one direct step could make the reconstruction cleaner and better at capturing motion.

Technical view

The method poses backbone recovery as indirect shape matching: an atomic point-cloud template is deformed so its simulated tomographic projections (through the cryo-EM imaging operator) match observed single-particle images, bypassing intermediate 3D potential-map reconstruction. The deformation is driven by a gradient flow on a Lie group, derived first in a general geometric setting then specialized to the SPA forward model. On synthetic data it recovers single- and multichain proteins and captures conformational transitions. A practitioner could extend it toward heterogeneous/continuous conformational analysis and, eventually, experimental data by plugging in realistic CTF and noise models into the projection operator.

arXiv · q-bio.QMBuildable

A Leakage-Proof Benchmark and Conformal Selective Triage for Electrohysterogram-Based Preterm Birth Prediction

An AI that predicted premature birth almost too well — because it was secretly cheating.

Doctors can pick up tiny electrical signals from a pregnant woman's abdomen (called electrohysterography, or EHG) that reflect the uterus's muscle activity, and researchers have tried using these signals to predict preterm birth. The catch: many past studies let recordings from the same patient show up in both the 'training' and 'testing' data, which is like letting a student see the exam answers beforehand — it makes the AI look smarter than it is. This paper builds a fairer test where each patient's data stays entirely on one side, and adds a system that flags cases the model is unsure about instead of forcing a guess. The goal is a more honest, trustworthy tool for identifying real preterm-birth risk.

Technical view

The authors formalize segment-level vs. patient-independent (record-grouped) validation on the Term-Preterm EHG Database (300 records, 38 preterm) and benchmark a 92-feature elastic-net logistic model under record-grouped nested cross-validation, with preprocessing, Platt calibration, and conformal estimation strictly confined to training folds. They further implement class-conditional conformal selective prediction to abstain on low-confidence cases rather than force uncertain classifications. This establishes a leakage-proof baseline other EHG classifiers can be benchmarked against using the same evaluation protocol.

arXiv · eess.SPBuildable

Empirical mode decomposition and interpretable machine learning for preterm birth classification from electrohysterography

Splitting a pregnant belly's electrical hum into layers to hunt for early-birth warning signs.

EHG signals from the abdomen carry a mix of overlapping electrical rhythms from the uterus, much like a song with multiple instruments playing at once. This study uses a technique called empirical mode decomposition to separate that mixed signal into simpler layers, then tests which layer (and which features extracted from it) best distinguishes women who deliver preterm from those who deliver at term. They also compare using expert-picked recording snippets versus plain fixed time windows, to see which gives more reliable results. The idea is to find a signal-processing recipe that's both accurate and doesn't accidentally cheat by mixing the same patient's data across training and testing.

Technical view

Using 26 recordings (13 preterm, 13 term) from the public TPEHGT dataset, the authors decompose EHG signals via empirical mode decomposition into intrinsic mode functions (IMFs) and extract 14 features per channel across 3 channels, comparing annotated intervals against non-overlapping 3-minute fixed windows. Nine classifiers are evaluated with repeated 5-fold recording-grouped cross-validation and recording-level aggregation to avoid the leakage problem noted in [1]. IMF1 (the highest-frequency component) gave the strongest mean classification performance, suggesting fast oscillatory content carries the most discriminative signal for term/preterm classification.

arXiv · q-bio.QMBuildable

DMT-Dens: Density-preserving manifold visualization for biological data

A cell-map tool that doesn't squash crowded neighborhoods into looking empty.

When scientists visualize thousands of individual cells based on their gene activity, they use 2D maps (like UMAP) to see clusters of similar cell types. But these maps often distort how 'crowded' or 'sparse' different regions really are, which matters when you're trying to spot rare or in-between cell states. DMT-Dens is a new mapping method, built on a transformer-based neural network, that specifically preserves this density information alongside the usual neighborhood structure. It works by making sure that how tightly packed points are in the original high-dimensional data matches how tightly packed they appear in the final 2D picture. This gives biologists a more trustworthy visual to spot rare or transitional cell populations.

Technical view

DMT-Dens is a parametric manifold-visualization method using a latent-token Transformer encoder that combines rank-based manifold alignment with hard-pair aggregation for neighborhood preservation. Its key addition is a density-preservation loss based on the Pearson correlation between k-nearest-neighbor log-radius estimates computed in the original high-dimensional space versus the 2D embedding space. Benchmarks show strong density fidelity on biological single-cell datasets, making it a drop-in alternative to UMAP/t-SNE when density-aware interpretation (e.g., detecting rare or transitional populations) matters, and being parametric, it can embed new/unseen data without retraining.

arXiv · q-bio.QMBuildable

Leveraging generative hallucination and biophysics-informed modeling for unified biomolecular sequence-structure co-design

An AI that designs new molecules by exploring branching 'what-if' possibilities like a chess engine.

Designing new biomolecules — proteins, but also trickier targets like DNA and RNA — that bind to a specific partner is central to drugs and biotech, but there's far less training data for DNA/RNA than proteins. MCTH tackles this by using existing AI models that predict 3D shapes from sequences (and vice versa) as building blocks, then uses a search strategy borrowed from game-playing AI (Monte Carlo Tree Search, the technique behind AlphaGo) to explore many candidate designs and spend its computing budget on the most promising ones. It factors in how confident the underlying models are and whether multiple predictors agree, and can optionally steer designs toward specific physical properties. This offers a way to design new molecule pairs without needing to retrain the underlying AI models.

Technical view

MCTH (Monte Carlo Tree Hallucination) is an inference-only framework that frames all-atom sequence-structure co-design as uncertainty-aware planning: it treats pretrained folding and inverse-folding models as frozen black-box operators, generating 'hallucinated' candidate states, and uses Monte Carlo Tree Search to allocate a fixed inference budget across competing design trajectories. Node selection incorporates model confidence/uncertainty and cross-expert consensus/disagreement when multiple predictors are available, with optional biophysical constraints folded into the same decision loop. Because it requires no retraining, practitioners can plug in any folding/inverse-folding model pair and apply it to non-protein modalities like DNA/RNA where labeled complex data is scarce.

arXiv · q-bio.QMBuildable

scDNM-VAE enables directly inspectable deep clustering of single-cell RNA-seq data through signed dendritic gating

A cell-clustering AI you can actually peek inside to see why it made each call.

When AI groups similar cells together from gene-expression data (single-cell RNA sequencing), it usually works like a black box — you get clusters but can't easily see the reasoning. scDNM-VAE is a new model inspired by how brain neurons process signals through branching dendrites, where each 'gate' has a clear direction, strength, and threshold for how it responds to the data. Because these gates are simple and explicit rather than buried in an opaque network, researchers can directly read off why a cell was assigned to a given cluster, without needing a separate explanation tool bolted on afterward. Tested on immune, brain, heart, and blood-stem-cell data, it holds its own against standard methods while being more transparent.

Technical view

scDNM-VAE pairs a variational autoencoder with a dendritic-neuron-inspired clustering head where cluster assignments are governed by learnable signed synaptic weights and thresholds: weight sign sets gate response direction, magnitude sets steepness, and the weight-threshold pair sets the transition location in latent space. This makes the trained clustering function directly inspectable without post-hoc explainability methods (e.g., SHAP/LIME analogs). It's benchmarked against scVI+KMeans and an MLP-DEC ablation across four datasets spanning immune, cortical, cardiac, and hematopoietic cells, offering a template for building interpretable-by-construction deep clustering models in other domains.

arXiv · nlin.AOConceptual

Phase-based spatial ordinal patterns for characterizing oscillatory dynamics

Reading the 'wave shapes' in brain or network rhythms like a fingerprint of their spatial pattern.

Many systems — from groups of neurons firing together to engineered oscillator networks — form visible spatial patterns as they pulse in sync or drift out of sync, and these patterns can shift suddenly and briefly (transient dynamics), which is hard to catch. This paper introduces a way to describe those patterns by looking at the timing (phase) of oscillations at nearby points and ranking their relative order, rather than looking at how strong the signal is (amplitude). From this ranking, they compute a single number — a kind of 'diversity score' — that rises when many different spatial patterns are present and can flag the moment a system briefly switches behavior. It's a general lens for studying rhythmic, spatially-spread-out systems, from brains to power grids.

Technical view

The method extends ordinal-pattern symbolic analysis to the spatial domain, operating directly on instantaneous phase rather than amplitude, with extra symbols added to handle near-equal phase values across neighboring points. This yields a symbolic representation encoding local spatial ordering that simultaneously captures phase gradients and synchronized clusters, from which a spatial permutation entropy is defined to quantify pattern diversity at each timepoint. The entropy time series enables detection of transient dynamics and regime shifts in oscillatory systems, giving practitioners a computationally light, model-agnostic diagnostic applicable to neural recordings or engineered oscillator networks.

arXiv · q-bio.PEConceptual

Dormancy stabilizes non-transitive competitive dynamics

Hitting 'pause' lets rock-paper-scissors species dodge the random extinctions that would otherwise wipe them out.

In ecosystems where species compete in a rock-paper-scissors style loop (each type beats one and loses to another), random population swings can accidentally wipe out a type entirely, collapsing the diversity even though no species is actually superior. Scientists knew that physical space — separate patches acting as refuges — can protect against this. This paper asks whether something similar can happen in time instead of space: specifically, whether organisms going dormant (like seeds lying inactive in soil, or bacteria entering a resting state) can act as a 'time refuge' that keeps lineages alive through unlucky stretches. Using a mathematical population model, they show dormancy indeed prevents this random collapse, offering a new explanation for how competitive diversity persists even in well-mixed, unstructured populations.

Technical view

The authors build a discrete-time Wright-Fisher population-genetic model that combines generalized seed banks with frequency-dependent (non-transitive, rock-paper-scissors-like) interactions, where an individual's type can be inherited from potential parents sampled across multiple past generations rather than only the immediately preceding one. This dormancy mechanism acts analogously to spatial structure, buffering lineages against interaction-driven stochastic fluctuations that would otherwise drive the system to fixation/extinction in a standard well-mixed model. The framework provides a tractable population-genetics tool for studying how temporal refuges (dormancy, seed banks) stabilize coexistence, applicable to microbial, plant seed-bank, or other systems with dormant life stages.

arXiv · math.APConceptual

Complete characterization of the sign of the wave speed in the symmetric Lotka-Volterra system under strong competition

A math proof settling which of two competing species wins the turf war along their border.

Imagine two species competing fiercely for the same space, spreading out and bumping into each other along a moving boundary — like two colors of mold racing across a petri dish. Mathematically, this is modeled with equations (Lotka-Volterra competition-diffusion) that predict a wave-like front between the two territories, and the key question is which species pushes the front forward and claims more ground. This paper proves, in full generality for the case where both species have identical competitive strength, that the species which spreads out (diffuses) faster always wins and expands its territory — except in the special case where both spread at exactly the same rate, where neither wins. It's a clean, complete answer to a question ecologists and mathematicians have long puzzled over.

Technical view

For the symmetric two-species Lotka-Volterra competition-diffusion system under strong competition (competition intensity >1), the authors fully characterize the sign of the unique bistable traveling front's speed: for every diffusion ratio d≠1, the front always expands the territory of the faster-diffusing species, with zero speed exactly at d=1. The key technical lemma is that no monotone standing front can exist when the two diffusion rates differ, which combined with continuity of wave speed in the model parameters yields the sign result; they also establish smooth dependence of the wave speed and front profile on parameters. This closes an open question in reaction-diffusion theory and gives a rigorous basis for predicting invasion outcomes in reaction-diffusion competition models from diffusion rates alone.

arXiv · math.PRConceptual

Order-Sensitive Fast-Synapse Limits in Sparse Excitatory-Inhibitory Threshold-Reset Networks

In simulated brain circuits, the order tiny signals arrive—not just their size—can flip whether a neuron fires.

This is a math study of simplified 'spiking' neuron networks, where neurons fire once their input crosses a threshold and then reset. Neuroscientists often simplify fast synaptic signals by shrinking their timing down to an instant, assuming only the total amount of excitation and inhibition matters, not the precise sequence they arrive in. The authors build two toy networks where the excitatory and inhibitory nudges shrink to the same instantaneous size in the limit, but arrive in opposite order — excitation-then-inhibition versus inhibition-then-excitation — and show a target neuron fires in one case but not the other. It matters because it exposes a hidden flaw in standard mathematical shortcuts used to model fast brain circuits: they can quietly give the wrong answer about whether a neuron actually fires.

Technical view

Uses a causal event-driven protocol with clamped refractoriness and smooth positive-delay kernels to construct two families of signed synaptic measures that converge weakly to δ₀ while their microscopic arrival order is reversed. The target neuron's firing condition, x+a−b<θ≤x+a, depends on this order rather than on the limiting measure alone, and the effect is shown to be robust to perturbations of state, pulse mass, and drift, persisting on sparse Dale-compatible random block graphs with q_N→∞, q_N/N→0. This implies that naive fast-synapse (instantaneous) limits of E/I spiking network models can be discontinuous or ill-posed, which matters for anyone deriving mean-field or diffusion approximations of spiking networks.

arXiv · cs.LGBuildable

PertMind: Eliciting Emergent Biological Reasoning in LLM via Reinforcement Learning on Cellular Perturbation Data

Training an AI on real cell-biology experiments—not textbooks—makes it reason like a biologist.

PertMind trains a large language model using actual lab measurements of how genes respond when a cell is perturbed, say by knocking out a gene or applying a drug. Instead of relying on expensive human-written explanations, it treats the measured outcomes as a reward signal in reinforcement learning, like a game score telling the model whether its prediction was right. The model starts with a supervised warm-up on trusted example reasoning, then improves through feedback scored at the level of individual genes, whole pathways, and answer formatting. Because it learned by predicting how perturbations play out, it also got better — without any extra training — at related tasks like figuring out what perturbation caused an outcome or picking the most promising experiments, suggesting it absorbed real biological intuition rather than memorized facts.

Technical view

PertMind combines supervised initialization on trusted reasoning trajectories with RL using multi-level rewards (gene-, pathway-, and format-level) computed directly from cellular perturbation atlas measurements, training only on forward perturbation-response prediction. It reports improved generalization to unseen cellular contexts plus zero-shot transfer to reverse perturbation identification, double-perturbation reasoning, phenotypic-screen prioritization, and biological-process interpretation without task-specific fine-tuning. This is a reusable template for converting large biological measurement atlases into RL environments with computable, non-human-curated rewards for post-training domain LLMs.

arXiv · q-bio.QMConceptual

tSymPerturb converts longitudinal symptom networks into time-indexed intervention strategies

A math toolkit tells therapists exactly which symptom to target, when, and how hard to push.

Psychologists increasingly model mental-health symptoms — sleep trouble, sadness, anxiety — as a network where each symptom can influence others over time. Knowing that symptom A tends to predict symptom B later doesn't tell a clinician what to actually do: how much to shift A, or whether nudging it will meaningfully change B down the line. tSymPerturb turns these descriptive prediction networks into an actual playbook, defining ways to simulate 'turning down' a symptom, blocking the pathway between two specific symptoms, and testing strategies like dosage, combining interventions, and choosing the best order to apply them. A core equation spells out exactly how much a later symptom will shift based on how much an earlier one was changed, turning a correlational map into something closer to a testable treatment-design tool.

Technical view

tSymPerturb extends the SymPerturb formalism to cross-lagged panel networks (CLPNs), separating source-state operators (virtual knockout/knockdown), transition operators (directed-edge or source-node communication blocking), and strategy procedures (dosage-response, combination analysis, sequence optimization). For a two-wave linear CLPN, the propagation identity Δμ₂ = B(μ₁ − μ₁*) makes explicit how a perturbation at wave 1 propagates through transition matrix B to wave-2 outcomes, and the paper derives falsifiable predictions — e.g., exact linear dose-response under a fixed intervention — that researchers can test against real panel data. This gives anyone with longitudinal CLPN data a route from correlational symptom networks to simulate-able intervention design.

arXiv · q-bio.NCBuildable

Continual-learning rules shape representational drift

How an AI avoids forgetting old skills quietly determines which of its internal memories drift over time.

Brains and AI systems both face the same puzzle: how do you learn new things without erasing what you already know, and does the trick you use to protect old memories change how those memories subtly shift over time? Researchers trained artificial networks — both image-classifiers and networks handling sequences of cognitive tasks — using different anti-forgetting tricks, most notably 'replay,' where the system periodically re-practices old material while learning new material. They tracked how the network's internal representation of the same fixed test inputs changed session after session, like repeated brain scans of the same thought. Replay kept performance intact, but the internal representations still drifted, and not randomly: deeper, more detailed processing stages wandered the most while coarse category information stayed stable — mirroring drift patterns seen in real animal brains, which suggests it's a natural byproduct of balancing stability and new learning.

Technical view

CNNs were trained on sequential image-classification tasks and RNNs on sequences of cognitive tasks under various continual-learning regularizers, with experience replay tracked in detail, while monitoring drift in fixed-probe representations across intervening tasks. Replay prevented catastrophic forgetting in both architectures, but representational drift still accumulated monotonically with the number of intervening tasks and was structured: later visual-processing layers and RNN temporal tuning drifted more than coarse class organization or task-relevant readout directions. This gives a testable, mechanistic link between specific continual-learning algorithms and drift signatures, letting neuroscientists compare recorded drift statistics in animals against model predictions to infer which mechanism better explains cortical drift.

arXiv · cs.LGBuildable

Multi-Feature Riemannian Hypergraph for Online Test-Time Adaptation of Motor Imagery Brain-Computer Interface

A brain-computer interface keeps working day after day by learning geometry, not just raw brain data.

Motor imagery brain-computer interfaces let a device guess what movement someone is imagining just from their brainwaves, potentially letting paralyzed patients control a wheelchair or robotic arm by thought. Two big problems plague real clinical use: brain signals look different day to day, and the system must adapt on the fly without stopping to retrain. MRieHy tackles this by first mathematically aligning each day's brain-signal patterns onto a shared reference using Riemannian geometry, a technique that treats the signal's covariance structure like points on a curved surface, then builds two web-like 'hypergraphs' connecting similar patterns — one from this geometric similarity, one from learned features — so the system can borrow strength from many related samples at once. Combining both lets the interface keep recognizing imagined movements accurately even as raw signals drift across days, tackling a major obstacle to everyday BCI use.

Technical view

MRieHy computes Riemannian means of EEG covariance matrices across cross-day training sessions to align multi-day distributions, then builds two complementary hypergraphs — one over covariance matrices using Riemannian distance, another over deep feature embeddings — to capture higher-order sample relationships during online test-time adaptation. This combines Riemannian-geometry domain alignment, standard in EEG transfer learning, with hypergraph message passing for higher-order (beyond pairwise) relations, applied specifically to the streaming/online setting rather than offline recalibration. Practitioners building MI-BCI pipelines could adopt the dual-hypergraph fusion as a drop-in module for test-time adaptation atop existing Riemannian-alignment baselines.

arXiv · q-bio.NCConceptual

A Control-Theoretic Formulation of Global Workspace Theory

A control-theory equation tries to pin down exactly which brain circuit makes things conscious.

Global Workspace Theory is a popular idea in consciousness science: the brain becomes aware of something when that information gets broadcast widely from a central hub to the rest of the brain. The problem is nobody has a precise, checkable definition of what counts as that hub. This paper borrows tools from control theory — the math engineers use to analyze how systems like thermostats or autopilots respond to and influence their surroundings — to define the hub as a subnetwork that can be driven by the rest of the brain, can in turn drive the rest of the brain, and has internal 'modes' linking the two directions in a distinctive way. By turning 'receives input,' 'sends output,' and 'transforms information' into precise mathematical properties, the authors create a testable signature that could, in principle, be checked against real or simulated brain circuits to see whether a candidate region really behaves like a global workspace.

Technical view

The Global Mediation Workspace (GMW) formalizes a candidate global-workspace subnetwork as an open dynamical system embedded in a larger network, using reachability (a Gramian capturing how external inputs drive the subnetwork), observability (how subnetwork states affect the rest of the network), and a boundary Hankel operator that identifies the internal modes coupling input-driven and output-driving dynamics. This gives a quantitative, control-theoretic alternative to informal 'broadcasting' language in Global Workspace Theory. It could be applied to whole-brain or large-scale RNN models by computing Gramians/Hankel singular values from simulated or empirical connectivity to test which subnetworks satisfy the GMW signature.

arXiv · q-bio.QMBuildable

Characterising cardiac tissue properties with graph neural networks

AI reads scattered heart sensors and pinpoints the exact scarred tissue causing dangerous heartbeats.

When doctors treat irregular heartbeats with a procedure called ablation, they need to find the exact patch of heart tissue causing the problem, but they can only take readings from a limited, sparse set of points inside the heart. This project trains a graph neural network — an AI good at reasoning over networks of connected points — on realistic simulated heart-signal data to spot suspicious regions, like scarred tissue, unusually fast-firing areas, or overly excitable tissue, all of which can trigger a dangerous rhythm called premature ventricular complexes. The model detected these regions with very high accuracy on simulated flat hearts, and with just a little extra fine-tuning it generalized to curved, more realistic heart shapes — a promising step toward guiding cardiologists to the right ablation target using fewer invasive measurements.

Technical view

The authors trained a GNN on synthetic electrogram data over 2D flat surfaces to classify localized regions of interest for PVC ablation, achieving average precision of 0.96 (fibrosis), 0.97 (rapid depolarization), and 0.95 (high excitability) from sparse intracardiac sampling. The model transfers to curved 2D surfaces via few-shot fine-tuning, indicating the learned graph representation generalizes beyond flat training geometry — a step toward geometry-agnostic clinical deployment. Real clinical use would still require validation on real patient electrograms and full 3D cardiac anatomy rather than synthetic flat-surface data.

arXiv · q-bio.QMBuildable

The Little Scientist: LLM Agent-Driven Discovery via the Scientific Method

An AI agent invents, tests, and revises its own scientific theories to design better algorithms.

This project asks: what if you gave an AI agent the actual scientific method — form a hypothesis, build it, test it, learn from results, repeat — and set it loose on inventing better algorithms? 'The Little Scientist' has a 'Scientist' AI agent work inside a testing environment that runs its code and reports detailed feedback on each case, much like a lab assistant handing back experiment results. When the Scientist gets stuck improving its own ideas, a second AI called the 'Kuhn agent,' named after the philosopher who coined 'paradigm shift,' steps in and throws it a wildly different idea borrowed from an unrelated field, forcing it to explore a totally different approach instead of endlessly tweaking the same one. The idea is that automated discovery needs deliberate disruption, not just iteration, to escape dead ends — mirroring how real scientific breakthroughs often come from outside conventional thinking.

Technical view

The Little Scientist implements an iterative loop where a Scientist LLM agent proposes and implements algorithm designs, evaluated by a benchmark environment returning structured per-instance diagnostics; on detecting a performance plateau, a Kuhn agent injects a cross-disciplinary 'paradigm-shifting' conjecture to redirect search away from the local optimum in the LLM's latent solution space. This is effectively an LLM-agent-driven automated algorithm design system with a built-in exploration/exploitation controller triggered by stagnation detection, demonstrated on two problems requiring different reasoning modes. Practitioners building LLM-agent AutoML or algorithm-search pipelines could adopt the plateau-detection plus cross-domain-analogy-injection pattern as a general escape-local-optima mechanism.

arXiv · cs.LGBuildable

Population Structure Analysis of an Inbred Population using Quantitative Shape Phenotyping from Stereo Retinal Photographs

AI reads the 3D shape of your eye's optic nerve to trace your ancestry.

Researchers studied a small, tight-knit community on Norfolk Island in the Pacific, many of whom descend from the Bounty mutineers, by looking at photos of the back of their eyes instead of their DNA. They took two-angle (stereo) photos of the optic nerve head — the spot where the eye connects to the brain — and computationally rebuilt its 3D shape, since that shape is partly inherited. A deep neural network then broke this shape down into layers of detailed features, and the most genetically telling features were used to sort people into ancestry-related clusters. The point is to show that a cheap eye photo can reveal population ancestry patterns that normally require expensive genetic testing.

Technical view

The pipeline reconstructs 3D optic nerve head (ONH) morphology from stereo fundus photographs via multi-scale stereo matching, then applies a self-taught deep learning model to extract hierarchical shape features at multiple scales. Features are ranked by discriminant power for distinguishing Bounty-descendant vs. non-descendant subpopulations within 781 Norfolk Island individuals, with performance validated via stratified cross-validation. Selected features feed hierarchical k-means-style clustering (k=2–7) to estimate admixture-like membership fractions, offering a low-cost phenotypic proxy for genetic population structure in imaging-based epidemiology.

arXiv · q-bio.NCConceptual

The effect of the excitatory feedback in anticipated synchronization and phase bistability regimes in neuronal populations

When brain regions talk back, a strange trick where the follower predicts the leader can break down.

In the brain, two connected regions can sync their rhythms, and oddly, sometimes the 'receiving' region seems to anticipate the 'sending' region rather than lag behind it — a bit like a dance partner predicting your next move. Scientists usually study this using a one-way connection, but real brain areas send signals back and forth. This paper adds that return signal (excitatory feedback) to computer models of two connected neuron populations and watches what happens to the anticipation effect and to a related phenomenon where the system can flip between two synchronization states. Understanding this helps explain confusing timing patterns seen in real brain recordings, where it's not obvious which region is 'leading.'

Technical view

The study extends unidirectional cortical-population models exhibiting anticipated synchronization (AS, negative phase lag from receiver having faster intrinsic dynamics) by adding excitatory feedback from receiver to sender, forming a bidirectional motif. Using coupled neural-mass-type oscillator models, the authors characterize how feedback strength modulates the existence and stability of AS versus delayed synchronization (DS), as well as the bistable regime between them. Results clarify how bidirectional cortical coupling — closer to physiological reality than the previously studied unidirectional case — shapes phase-lag statistics observed in electrophysiology, informing interpretation of lead-lag relationships in real inter-areal recordings.

arXiv · q-bio.TOBuildable

Head Impact Characterization and Cellular Response of a Live-neuron cell-integrated Biomechanical Full-body Surrogate Model

Scientists dropped a crash-test dummy with live brain cells inside its head to watch neurons react to impact.

To understand what actually happens to brain cells during a head injury, researchers built a crash-test-dummy-like full-body model with real, living neurons embedded inside its head in small dishes. They dropped the dummy from a seated position at different angles to simulate falls, then measured the forces on the head with accelerometers while simultaneously tracking how the living cells responded to the jolt. A separate computer model of human muscles and bones was used to double-check the fall's physics. The goal is to directly connect the mechanical violence of an impact to the biological damage it causes at the cellular level, which could improve helmet design and injury thresholds.

Technical view

The framework integrates a commercial anthropomorphic surrogate with three stacked Petri dishes of live SH-SY5Y neuroblastoma cells inside the head, instrumented with six accelerometers (three head-surface, three in-series with the cell stacks) to capture impact kinematics during controlled seated falls at 30°, 60°, and 90° release angles. An OpenSim musculoskeletal model runs in parallel to reproduce fall biomechanics, enabling correlation of measured head acceleration/deformation with observed cellular response. This links macro-scale biomechanical impact metrics directly to micro-scale cellular outcomes, offering a validation platform for injury-threshold and protective-equipment research.

arXiv · q-bio.PEConceptual

Stochastic Pharmacokinetic Escape: A Field-Theoretic Approach to Fluctuation-Induced Tumor Relapse

Math shows why chemo that looks like it 'cures' a tumor on paper can still let it come back.

Standard cancer-treatment math assumes that if you give a high enough dose of chemotherapy, the tumor's cell count smoothly goes to exactly zero — a clean cure. This paper argues that's an illusion caused by ignoring randomness: real cell populations are small, discrete, and noisy, especially in tiny hidden pockets of tumor cells shielded from the immune system. Using advanced physics techniques for modeling random fluctuations (originally built for other 'noisy population' problems), the authors show these tiny surviving clusters have a genuinely nonzero chance of regrowing even after treatment that looks perfect on average. This matters because it offers a mathematical reason why cancers relapse even after seemingly successful chemo, and could inform how doses are scheduled.

Technical view

The authors build a nonequilibrium stochastic PK-PD field theory coupling a two-compartment pharmacokinetic model to a stochastic tumor-immune sector, formalized in the Doi-Peliti operator formalism and mapped to multiplicative Langevin equations via the Martin-Siggia-Rose/Janssen-De Dominicis path-integral approach. In immune-depleted sanctuary sites the dynamics reduce to a time-dependent Feller diffusion process, and the corresponding Fokker-Planck equation yields a closed-form survival functional showing that demographic noise gives micro-clusters a strictly positive relapse probability even under deterministic-cure-predicting high-dose bolus chemotherapy. This provides an analytical, testable framework for reassessing dosing strategies (e.g., metronomic vs. bolus) against fluctuation-driven relapse risk.

arXiv · cond-mat.softConceptual

Growth-Induced Transitions in Viscoelastic Matter

Growing tissue isn't just solid or liquid — push it fast enough and it does something entirely new.

Living tissues like tumors or biofilms grow, and as they grow they also respond to stress like a mix of a solid (springy, elastic) and a liquid (flowing, viscous) — think of Silly Putty. This paper asks what happens when the speed of growth starts to compete with how quickly the material relaxes stress internally. Using a simple test case — a growing elastic beam — the authors find that when growth and relaxation happen at similar speeds, the material doesn't just gradually shift between solid-like and liquid-like behavior; it undergoes a genuinely new kind of transition with its own distinct dynamics. This reshapes how we should think about the mechanics of anything that grows while also being squishy, from tumors to bacterial colonies.

Technical view

The paper analyzes proliferating viscoelastic matter where growth rate g competes with the material's viscoelastic relaxation time τ, showing the combined dimensionless parameter gτ governs a qualitative transition rather than a smooth interpolation between the purely viscous (gτ→0) and purely elastic (gτ→∞) limits. Using a growing elastic beam as the canonical test system, they identify new dynamical regimes emerging at intermediate gτ that are absent from either limiting theory. This establishes a general theoretical lens — applicable to biofilms, tumors, and other proliferating soft matter — for predicting when growth-induced mechanical instabilities (e.g., buckling, morphogenesis) will deviate from standard elastic or viscous growth models.

arXiv · q-bio.PEConceptual

Phylogeny-based metrics of biodiversity: concepts and methods

A review of the math tools conservationists use to save the whole 'tree of life,' not just species counts.

When deciding which species to prioritize for conservation, just counting species can miss the bigger picture — losing one weird, evolutionarily unique species (like a platypus) is a bigger loss to life's diversity than losing one of many similar frog species. Since the early 1990s, scientists have built mathematical tools that measure diversity using the 'tree of life,' the branching family tree connecting all species, so that older, more distinct branches count for more. This chapter walks through how those tools evolved, especially a widely used one called phylogenetic diversity and its descendants (like EDGE and EDGE2), which rank individual species by how much unique evolutionary history they'd take with them if they went extinct. It matters because it gives conservationists a more principled way to decide where limited funding and effort should go.

Technical view

The chapter reviews the methodological lineage of phylogenetic diversity (PD) metrics, starting from Faith's PD (sum of edge lengths in a rooted phylogenetic subtree) through species-specific prioritization indices like EDGE (Evolutionarily Distinct and Globally Endangered) and its refinement EDGE2, which quantifies expected marginal PD loss per species under extinction risk. It surveys the mathematical formalization and extensions of this framework for use in conservation prioritization, providing a technical primer for practitioners implementing PD-based ranking in biodiversity assessment or reserve-design software.

arXiv · q-bio.NCConceptual

Valhalla: A Layered Knowledge-State and Service-Governance Framework for Long-Term Scientific Knowledge Work

A new filing system lets AI research assistants share and organize scientific knowledge across an entire team.

AI assistants that help with science increasingly rely on memory systems — knowledge graphs — to keep track of files, ideas, and results over long projects. The problem is that most of these systems are built around one person's personal organizing habits, making it hard for a whole team to share and reorganize that knowledge together. Valhalla is a proposed framework that replaces the usual flat, tangled web of notes with organized layers: separate levels for raw files, extracted resources, defined concepts (entities), the relationships between them, and the overall graph. This layered structure is meant to make long-term scientific knowledge easier to trust, share, and rebuild as a research team's understanding evolves.

Technical view

Valhalla introduces a five-layer File-Resource-Entity-Relationship-Graph (FREG) architecture as an alternative to conventional flat, node-centric knowledge graphs used for LLM-agent long-term memory in scientific workflows. Files and Resources preserve source identity and provenance, while Entities and Relationships are built as stable semantic abstractions on top, layered under a governing Graph service — aiming to decouple knowledge structure from any single user's ad hoc organizational scheme and enable cross-user sharing, integration, and reorganization. This targets a concrete pain point in multi-agent/multi-user RAG and knowledge-management systems: provenance-preserving, reorganizable shared memory, relevant to anyone building persistent knowledge infrastructure for LLM research agents.

arXiv · q-bio.NCBuildable

Phase- and amplitude-dependent control of synchronization in excitatory-inhibitory networks via pulsed stimulation

A single, well-timed electrical pulse can push a network of neurons into or out of sync — if you hit it right.

Groups of neurons often fire in rhythmic waves, and scientists want to know how a brief external nudge — like a pulse of current, similar to what a stimulation device might deliver — changes that rhythm. The usual tool for this, the phase response curve, only tracks how the pulse shifts the timing of the rhythm, but it misses how the pulse also changes the strength or intensity of the synchronized activity. This paper studies a simulated network of excitatory and inhibitory neurons and tracks both the timing shift and the intensity shift caused by pulses delivered at different points in the rhythm. They find that the exact same pulse can make the network more synchronized or less synchronized purely depending on when in the cycle it arrives, which matters a lot for designing brain-stimulation therapies that aim to calm or boost rhythmic brain activity.

Technical view

The authors simulate a balanced excitatory-inhibitory network of exponential integrate-and-fire (EIF) neurons receiving phase-targeted transient current pulses, and jointly compute the network phase response curve (nPRC), a novel network amplitude response curve (nARC), and resulting changes in population synchrony as functions of stimulus timing. They demonstrate that identical pulses can enhance or suppress synchronization depending solely on pulse phase, showing phase-resetting theory alone (standard PRC analysis) is insufficient to predict collective network response — amplitude effects are equally causal. This nPRC/nARC framework gives a quantitative, testable basis for designing closed-loop or phase-locked neurostimulation protocols aimed at modulating pathological or therapeutic network synchrony (e.g., in epilepsy or Parkinsonian oscillations).

arXiv · q-bio.NCConceptual

Synaptic delays modulate population phase and amplitude responses in oscillatory excitatory-inhibitory networks

How long it takes one neuron to nudge another secretly tunes whether brain rhythms wobble or snap back.

Brain cells talk to each other through synapses, and there's always a tiny delay before a signal from one neuron actually affects the next. This study looks at brainwave-like rhythms produced when excitable 'go' neurons and calming 'stop' neurons volley signals back and forth in a fast, well-known rhythm pattern seen in the cortex. The researchers poked these simulated networks with brief jolts and measured two things: whether the rhythm's timing got shifted (like a beat dropping early or late) and whether its strength changed. They found that the delay length itself — not just the network's wiring — determines how the whole population recovers from a disturbance, which matters for understanding how the brain keeps rhythms stable or lets them shift, something linked to attention, memory, and disorders like epilepsy.

Technical view

Using a conductance-based spiking network in the PING (pyramidal-interneuron gamma) regime, the authors systematically vary synaptic delay and apply brief perturbations to excitatory (E), inhibitory (I), or combined populations, then compute network phase response curves (nPRCs) and amplitude response curves (nARCs). This extends single-neuron PRC theory to a network observable, showing delay reshapes both timing-reset and amplitude-recovery dynamics of the collective oscillation, not just the E/I balance typically emphasized. Practitioners modeling gamma oscillations or building delay-coupled E-I mean-field models can use nPRC/nARC as a diagnostic to predict entrainment and desynchronization thresholds under stimulation (e.g., for closed-loop neurostimulation design).

arXiv · eess.SPBuildable

CORAL: A Modality Invariant Framework for Robust Vital Sign Rate Estimation Using Correloform Analysis

One math trick reads your heartbeat and breathing from almost any wearable sensor, cutting through noise automatically.

Doctors and wellness apps need to track heart rate and breathing rate from all sorts of sensors — chest straps, fingertip clips, hospital monitors — but each sensor type usually needs its own custom software to filter noise and find the rhythm. CORAL is a new general-purpose method that turns any repeating, wave-like signal into a special 2D picture (they call it a 'correloform') that highlights how periodic the signal is over time, using a classic statistical tool called autocorrelation. Because this picture-making process is built from math rather than tricks tailored to one device, the same system works across many sensor types without retooling, automatically figuring out which sensor channel is cleanest and flagging bad data. This matters because it could let hospitals and at-home health devices reliably measure vital signs no matter what hardware is attached.

Technical view

CORAL reintroduces Short-Time Autocorrelation Functions (STACFs) to construct a 'correloform' — a 2D time-varying representation of signal periodicity — as a modality-agnostic backbone for rate estimation from quasi-periodic biosignals (ECG, PPG, SCG, BioZ). Because the transform is derived from generic autocorrelation mathematics rather than modality-specific filtering/feature engineering, the same pipeline yields rate estimation, noise resilience, automatic channel selection, and signal quality indicators across signal types without retraining. This is relevant to anyone building multimodal vital-sign monitoring pipelines (hospital or wearable) who wants a single robust front-end rather than maintaining per-modality DSP chains; benchmarking across in-hospital and at-home conditions suggests it generalizes across population and noise regimes.

arXiv · q-bio.PEConceptual

Social Discounting Enables Fast and Reliable Collective Escape

Watching a calm neighbor not flee is itself proof there's no danger — and groups exploit that math.

When an animal senses a possible predator, it faces a tradeoff: react fast and you'll often be wrong, react carefully and you might react too late. This paper shows that animals in a group can beat that tradeoff by paying attention not just to neighbors who bolt in fear, but also to neighbors who stay calm — because calm neighbors are quietly telling you 'I don't see a threat either.' The researchers modeled each animal as gradually gathering evidence of danger until it crosses a mental tipping point to flee, then compared a 'naive' animal that only reacts to others fleeing (which gets trigger-happy in bigger groups) against a 'smart' Bayesian animal that also credits calm as evidence of safety. The smart strategy lets the whole group escape real threats faster and with fewer false alarms, and it explains why panic can either fizzle out or cascade explosively through a herd depending on whether real danger is present.

Technical view

The authors model each individual as a drift-diffusion evidence-accumulator with a flee threshold, then formalize social inference by treating both neighbor flight (positive evidence) and neighbor stillness (negative evidence) as informative signals, collapsing the optimal Bayesian weighting into a single 'social discounting rate' parameter interpolating between naive (flight-only) and full Bayesian responders. This yields closed-form expressions for group detection speed, false-alarm rate, and cascade branching ratios, showing the branching ratio stays subcritical under safety but becomes supercritical under real threat — explaining self-limiting vs. explosive alarm cascades as an emergent property of the discounting rate rather than added mechanism. This provides a tractable analytical framework (rather than pure simulation) that collective-behavior and neuroscience-of-decision-making researchers could extend to empirical flocking/herding datasets or robotic swarm alarm protocols.

arXiv · cs.LGBuildable

PathFinder: Joint Decompositions of Linked Multimodal Datasets

A new algorithm finds shared patterns across datasets that don't even overlap directly, just via chained connections.

Scientists often have several different datasets — say gene activity, patient scans, and drug response data — that they want to analyze together to spot shared patterns, but standard methods require every dataset to share some common dimension, like the same patients or same genes measured in each. PathFinder gets around this by allowing datasets to connect indirectly: if dataset A shares something with dataset B, and B shares something with C, PathFinder can chain those overlaps into one combined analysis even though A and C never directly overlap. It works like a matchmaking network, tracing valid 'paths' through the web of shared connections to build one unified low-dimensional summary of everything. This matters because real-world data collected across different species, labs, or measurement types rarely lines up neatly, so this method could unlock joint analysis that used to be impossible.

Technical view

PathFinder extends joint/linked low-rank matrix decomposition methods (e.g., joint NMF/PCA variants) to settings where the full dataset collection lacks a common shared dimension, but pairwise or subgroup-level shared dimensions form a connected graph. By identifying paths through this graph of matrix-to-matrix shared dimensions, PathFinder propagates factor constraints transitively to estimate a global joint decomposition, recovering common latent patterns across modalities, species, or scales without requiring a universal one-to-one sample/feature mapping. This is directly applicable to multi-omics or cross-species integration pipelines where researchers currently must discard datasets that don't share an axis with every other dataset; implementation would build on existing joint-NMF/CCA-style optimization with graph-based path constraints.

arXiv · q-bio.QMRunnable

MultiStructRNA: a Python package for multi-algorithm RNA secondary structure prediction, ensemble analysis, and visualization

One Python toolkit finally lets you run and compare every RNA-folding algorithm without rewriting your code each time.

RNA molecules fold into shapes (secondary structures) that determine how they work in cells and how well drugs can target them, and scientists use many different computer algorithms to predict these shapes. The problem is that each algorithm has its own quirky input format, output style, and lack of built-in visuals, making it a hassle to compare methods or combine their results. MultiStructRNA solves this by giving researchers one simple interface that can run many different prediction algorithms, translate all their outputs into a shared, consistent format, and generate clear visualizations, whether working in a notebook or a large automated pipeline. This matters because it turns a fragmented, error-prone process into something researchers can trust and reproduce, especially for large-scale RNA studies used in drug and vaccine design.

Technical view

MultiStructRNA is a Python package that wraps multiple RNA secondary structure prediction algorithms behind a unified high-level API, standardizing heterogeneous inputs/outputs into a consistent object schema and computing ensemble-aware consensus metrics across predictors. It's built for both interactive (Jupyter) and production (pipeline/scripted) use, supports high-throughput batch analysis, and includes native visualization for structure comparison. Bioinformaticians can use it as a drop-in orchestration layer to benchmark or ensemble existing folding tools (e.g., comparing thermodynamic vs. comparative methods) without writing custom format-conversion glue code for each one.

arXiv · eess.ASRunnable

A Parameter-Free Few-Shot Evaluation for Elephant Vocalisation Classification

With just a handful of examples per call, a dead-simple 'nearest average' method can classify elephant sounds.

Researchers who want computers to recognize different types of elephant calls usually train complex classifiers on lots of labeled examples, but labeling animal sounds is slow and expensive. This study instead asks a simpler question: if you only have a few labeled examples per call type, how well does the most basic possible method work — just averaging the 'fingerprint' (embedding) of each known call type and matching new sounds to whichever average is closest? They test this simple approach using several pre-trained sound-recognition AI models as the fingerprint-makers, across real elephant vocalization datasets, and repeat the test many times with randomly resampled small example sets to check reliability. This matters because if a simple, parameter-free method works nearly as well as complex trained classifiers, it could make wildlife bioacoustics research faster and more accessible with far less labeled data.

Technical view

The authors evaluate nearest-centroid ('prototypical') classification on frozen pretrained acoustic embeddings (Perch v1, Perch v2, HuBERT-base layer 2, and MFCC baselines) for elephant call classification, using an N-way k-shot episodic protocol with class prototypes computed as the mean of support-set embeddings and query assignment via nearest centroid in squared Euclidean distance. Evaluated across the Elephant Voices and LDC datasets with 100-resample bootstrapping over support sets, this establishes a parameter-free lower-bound baseline against which trained classifiers can be benchmarked as labeled data scales from few-shot to full. Bioacoustics practitioners can use this as a near-zero-cost baseline pipeline (no training loop, just embedding extraction + centroid distance) before investing in fine-tuned classifiers for new species or call types.

arXiv · q-bio.PEConceptual

Extinction drives emergent metastability in complex ecosystems

Random chance killing off rare species turns out to be what keeps whole ecosystems from collapsing.

Every species eventually goes extinct, but scientists have debated for decades whether having more species in an ecosystem makes it more or less stable overall. Most past models treated species survival as a fixed, deterministic outcome, ignoring the fact that small, rare populations can randomly die out just from bad luck (like a run of failed births), a phenomenon called demographic stochasticity. This paper builds a model that includes that randomness and finds it actually helps: random extinctions quickly weed out the weakest, lowest-population species, leaving behind a leaner ecosystem that's more stable overall. They also discover the pattern of how many species survive over time follows an unusual statistical shape with a 'heavy tail,' meaning occasional dramatic extinction events are more common than you'd expect. This reframes extinction not as pure loss but as a stabilizing pruning process in complex ecosystems.

Technical view

The authors incorporate demographic stochasticity into a rule-based large complex ecosystem model (extending classic random-matrix stability-diversity frameworks like May's), showing that stochastic extinction events preferentially prune low-abundance species and thereby increase the stability of the surviving community relative to deterministic-extinction baselines. They develop a bottom-up analytical theory characterizing extinction statistics and find the surviving-species fraction follows an anomalous heavy-tailed distribution rather than the exponential/Gaussian decay typically assumed. This offers theoretical ecologists a stochastic extension to classical stability-diversity theory and a testable statistical signature (heavy-tailed survival fraction) that could be checked against empirical community time-series or long-term ecological survey data.

arXiv · q-bio.GNConceptual

Ten simple rules for non-visual, reproducible and accessible bioinformatics

Ten practical rules to make gene-data science work without needing to see a single chart.

Modern biology research relies heavily on visual tools — colorful plots, heatmaps, and interactive charts — to make sense of huge datasets and decide what the results mean. But for blind and low-vision scientists using screen readers or braille displays, these visuals are often inaccessible, meaning the actual evidence behind a scientific decision is locked away in a format they can't use. This paper argues that fixing this accessibility gap has a bonus: it also makes research more reproducible for everyone, because both goals demand that every analytical decision be written down and explained in text rather than left implicit in a picture. The authors lay out ten concrete practices, like treating plots as documented decision records, using AI carefully to describe figures in words, and writing code and workflows that are text-first, so blind and sighted researchers alike can follow and verify the reasoning. This matters for making science genuinely open to more people and more trustworthy overall.

Technical view

The paper presents ten practical guidelines for non-visual bioinformatics aimed at researchers using screen readers, braille displays, or audio interfaces, framing visual analysis artifacts (QC plots, embeddings, heatmaps, genome browser tracks) as undocumented decision points that should instead be captured as explicit, text-based decision records. Recommendations span cautious use of AI-generated alt-text/figure descriptions, accessible computing environment setup, and text-first literate programming practices (e.g., structured markdown/notebook output over rendered-only graphics) that double as reproducibility documentation. Bioinformatics tool developers and lab leads can use this as a concrete checklist for auditing pipelines (e.g., ensuring every QC gate has a textual threshold/rationale, not just a plotted cutoff line) to improve both accessibility and audit-trail rigor.

arXiv · physics.soc-phConceptual

Absorbing phase transition in a queueing model of coupled adaptive agents

A model of "will you show up?" shows why shared plans can suddenly and permanently collapse.

This is a mathematical model of how people decide whether to do things together (like meeting a friend) or alone, when limited time is split between shared and private tasks. Each person weighs the value of the joint activity against the risk that the other might flake, so joining becomes a calculated bet rather than random chance. The researchers found that as this risk grows, the system doesn't drift gradually from "we do things together" to "everyone acts alone" — it snaps abruptly between the two, like a switch flipping. Once it falls into the solitary state it's very hard to recover from; someone has to stubbornly keep showing up alone for a while to pull the group back. This matters because it explains why social coordination in friendships, teams, or communities can break suddenly and stay broken.

Technical view

The authors extend the priority-queue model of human activity by letting agents endogenously set the priority of a joint task, discounted by their belief about a partner's participation probability, rather than drawing priorities from a fixed distribution. This strategic feedback introduces a phase transition between a "coupled" phase of sustained joint activity and an absorbing "solitary" phase, separated by a saddle-node bifurcation derived in closed form. The transition is discontinuous, and the solitary phase is absorbing so the system cannot spontaneously re-enter the coupled phase unless an agent unilaterally persists in the joint activity for roughly one memory-time, a cost the paper also quantifies. This gives a tractable analytical handle on hysteresis and irreversibility in coordination dynamics, applicable to modeling social-tie fragmentation or dyadic relationship breakdown.

arXiv · q-bio.PEConceptual

Body size predicts how long ant workers live - but not how they age or how they die from heat

Bigger ants live longer, but size doesn't explain how fast they age or survive deadly heat.

Scientists wanted to know whether a worker ant's body size determines not just how long it lives, but also how it declines with age and how well it survives extreme heat — three questions usually lumped together as one. They studied 18 Australian ant species, tracking thousands of individual ants across paired field and lab settings to measure survival over time. Bigger ants reliably lived longer, likely due to basic physiology, while colony size and temperature had little effect on that pattern. Surprisingly, how fast an ant ages had nothing to do with size — instead it tracked when the species is active during the day, aging fastest in species active in the morning. This matters because it shows lifespan, aging, and heat tolerance follow different biological rules, which is important as climate change raises heat stress on ecosystems.

Technical view

Using paired field-laboratory survival assays across 18 Australian ant species (2,363 cohort-day observations, 1,148 workers), the study decomposes mortality risk into three components: lifespan duration, senescence trajectory, and thermal vulnerability. Cox proportional-hazards models show body size significantly predicts duration (HR = 0.67, p = 0.002) independent of colony size or a size×temperature interaction (both non-significant), with only a weak size×foraging-rate interaction (LRT p = 0.014) suggesting mostly intrinsic physiological drivers. Senescence trajectory instead tracks circadian niche (Kruskal-Wallis p = 0.009, steepest in matinal/morning-active species) rather than body size, decoupling aging rate from both lifespan and thermal death risk. The dataset offers a rare multi-species, multi-axis mortality decomposition useful for comparative life-history and climate-vulnerability modeling in social insects.

arXiv · eess.IVBuildable

Test-Time Instance Selection for Improved Whole Slide Image Analysis

A plug-in trick lets pathology AI skip boring tissue patches and focus only on the telling ones.

When AI analyzes a giant medical slide image, chopped into thousands of tiny patches, most patches are uninformative filler tissue that wastes computing power and can dilute the signal from the handful that actually matter for diagnosis. The researchers built TTIS, a method that at the moment of analysis automatically selects the small set of patches best representing the tissue, without any extra training — it just plugs into existing AI models. It also examines the slide from multiple views to make the selection more reliable. This matters because it could make cancer-diagnosis AI faster and sharper without the cost of retraining models on new hospital data.

Technical view

TTIS is a training-free, plug-and-play framework for whole slide image (WSI) multiple instance learning (MIL) that performs instance selection purely at inference time, choosing a compact, representative patch subset instead of processing every extracted patch. It adds a multi-view ensemble strategy that aggregates distinct facets of tissue morphology to improve selection robustness. Because it requires no retraining, it drops directly into any existing MIL pipeline (e.g., ABMIL, TransMIL, CLAM-style architectures) as an inference-time module, making it immediately deployable on pretrained pathology models. The practical payoff is reduced inference redundancy and compute alongside potential accuracy gains from filtering non-informative patches.

arXiv · eess.IVBuildable

KHiM-Mamba: Injecting Pathology Knowledge into Mamba via Hidden-State Modulation for Whole Slide Image Analysis

Feeding pathology know-how into an AI's memory keeps it focused on the slide regions that matter.

This tackles the same giant-medical-image problem as before, but with Mamba, a newer AI architecture efficient at handling very long sequences of data like all the patches in a huge slide. The catch is that Mamba, looking only at raw visual features, can get distracted — irrelevant tissue crowds out the small but crucial diagnostic regions as it scans, diluting key evidence over the long "reading" process. The fix is to inject actual pathology knowledge into the AI's internal hidden memory state, nudging it to attend to what a pathologist would consider medically meaningful rather than just visually eye-catching. This matters because it could make AI slide analysis both faster, thanks to Mamba's efficiency, and more clinically trustworthy.

Technical view

KHiM-Mamba addresses a failure mode in selective state-space model (SSM/Mamba)-based MIL for WSIs, where purely vision-driven state updates misallocate attention across long patch sequences, causing the SSM's hidden state to accumulate task-irrelevant evidence and dilute diagnostically decisive regions. The proposed Knowledge-Aware Hidden-State Modulation architecture injects external pathology knowledge directly into the Mamba hidden state, biasing its selective update/readout dynamics toward clinically meaningful regions rather than relying solely on visual saliency. This preserves Mamba's core advantage — linear complexity over long sequences — while correcting its knowledge-blind selectivity, positioning it as an alternative to attention-based MIL aggregators for gigapixel WSI encoding. Practitioners on SSM-based pathology pipelines could adopt the hidden-state modulation mechanism as a general way to fuse domain priors into SSM state dynamics beyond WSI analysis.

arXiv · q-bio.PEConceptual

An Analytically Tractable Framework for Multi-Strain Epidemics: Resolving Algebraic Complexity to Map Oscillatory Dynamics

A clever simplification finally lets scientists write exact formulas for why epidemics wax and wane in cycles.

Diseases with multiple strains, like flu variants, often arrive in repeating waves rather than settling into steady state, but the math describing two interacting strains has been too messy to solve exactly. The researchers found a special, simplified version of the two-strain model — where immunity to one strain fully blocks second infections from it in an asymmetric way — that turns out to be exactly solvable with clean formulas. Using this, they precisely map when the disease settles into steady coexistence versus when it breaks into self-sustaining oscillating outbreaks, driven by how much more infectious "second-time" infections are. Simulations further reveal that small everyday fluctuations and big dramatic outbreak cycles can occur side by side in the same system. This matters because it gives epidemiologists exact mathematical tools, not just simulations, for understanding multi-strain disease cycles like flu or dengue.

Technical view

The paper identifies an analytically tractable subclass of two-strain SIR-type models with asymmetric cross-immunity, where immunity to one strain fully blocks secondary infections from that strain, eliminating the algebraic complexity that has historically blocked closed-form analysis of multi-strain coexistence. Within this class, the authors derive explicit expressions for coexistence equilibria and their stability, showing the relative transmission advantage of secondary infections is the key bifurcation parameter governing onset of limit-cycle oscillations, generalizing prior narrow near-threshold results to a broader regime. Numerical simulation of the full (non-simplified) system reveals a richer bifurcation landscape than the tractable subcase alone predicts, with small-amplitude local oscillations coexisting with large-amplitude recurrent outbreak cycles. This offers epidemiological modelers a reference analytical benchmark for validating numerical multi-strain models and for reasoning about strain-competition-driven oscillatory dynamics such as influenza subtype cycling or dengue serotype dynamics.

arXiv · q-bio.NCConceptual

Data-driven techniques for translational neuroscience and personalized neuro-health

A survey maps the toolbox of data methods racing to catch brain disease before it's too late.

Diseases like Alzheimer's and Parkinson's are usually diagnosed only after a lot of irreversible brain damage has occurred, so there's a big push to spot subtle warning signs earlier using brain scans and data analysis. This paper is a review that rounds up a wide range of modern data-driven methods, organized into four main categories, that researchers use to detect small, early, person-specific brain changes. Rather than inventing a new method, it maps the current landscape, showing how these diverse approaches all aim at the same goal: personalized models of brain health grounded in real biology and useful in the clinic. It also lays out what's still unsolved statistically, computationally, and clinically. This matters because it could help clinicians catch neurodegenerative disease earlier, when treatment might still make a difference.

Technical view

This review surveys data-driven methods for translational neuroscience and personalized neuro-health, organized around four methodological pillars spanning neuroimaging-based quantitative biomarker detection for early, individualized neurodegenerative change. Its throughline is convergence: despite methodological diversity, these approaches share the translational goal of producing personalized, mechanistically grounded, clinically actionable brain-health models for diseases like Alzheimer's and Parkinson's, currently diagnosed only after substantial irreversible neuronal loss. It closes by cataloging open statistical, computational, and clinical challenges, making it a useful orientation point for researchers deciding which methodological pillar — specific imaging biomarkers, modeling frameworks, etc. — to build on for early-detection or precision-neurology work. As a review, its value lies in synthesis rather than a novel result to replicate directly.

arXiv · q-bio.PEConceptual

Shared environmental risk selects asymmetric inheritance of a protective reserve

A math model shows why cells sometimes deliberately split their emergency supplies unevenly between offspring.

When a cell divides, it must decide how to split a limited protective reserve — like a rainy-day fund against stress — between its two daughter cells. You'd assume splitting evenly is always safest, but this model shows that when environmental risk is shared between the daughters (say, they face the same conditions), giving unequal amounts to each can actually be better for long-term survival. Using a mathematical measure of long-run growth rate under randomness, the researchers find this unequal-split strategy suddenly becomes favored once environmental sharing crosses a threshold, happening as an abrupt jump to a specific asymmetry level rather than a gradual drift. This holds up even when other model details are changed. This matters for understanding why asymmetric division shows up in biology, such as in stem cells or bacteria.

Technical view

The paper models cell division as reserve-partitioning between mother and two daughters, where each fixed partition policy (parameterized by asymmetry α) generates a random demographic operator whose top Lyapunov exponent determines the population's long-term growth rate under environmental fluctuations. Environmental sharing between sibling lineages is coupled to the value of diversification: weak sharing favors symmetric (α≈0.5) inheritance, while sufficiently strong shared risk selects a distinct asymmetric branch near α≈0.2, with the transition occurring as a discontinuous jump between fitness-optimal branches rather than a continuous bifurcation. This asymmetric-optimal phase is robust to changes in the protection law and reserve turnover dynamics, though the exact transition boundary depends on protection nonlinearity and reserve memory — a Lyapunov-exponent framework that could be extended to test specific molecular reserve systems, such as damage or chaperone partitioning, against measured asymmetry ratios.

arXiv · cs.CLBuildable

Localize, Then Reason: Visual Latent Structural Reasoning for Molecular Properties and Edits

Teaching AI to zoom in on the right part of a molecule before reasoning speeds it up nearly 10x.

To predict a molecule's properties or suggest edits, an AI needs to know which specific parts of its structure matter chemically — but current models either get told which parts matter by humans or must guess properties from a full image without knowing where to look. VLSR teaches the AI to first spot the chemically important regions in a picture of the molecule on its own, then reason about how those regions affect its properties inside a compact internal workspace, rather than describing everything at length. It's like teaching someone to circle the important part of a diagram before explaining what it does, instead of narrating the whole picture. This localize-then-reason approach lets the AI process molecules 9.6 times faster than a comparable baseline. This matters for drug discovery and materials science, where AI needs to reason about chemical structures quickly and accurately at scale.

Technical view

VLSR (Visual Latent Structural Reasoning) is an end-to-end framework for molecular property and edit reasoning that jointly learns localization of chemically meaningful regions in a molecular image and property reasoning over those regions, rather than relying on externally provided motif annotations or reasoning directly and unstructured from full images. Its core "localize-then-reason" strategy first learns region localization, then performs property-effect reasoning in a compact latent workspace before decoding the final answer, avoiding verbose token-level chain-of-thought over raw pixels. Under matched inference settings, this design achieves 9.6x higher throughput than a comparison LLM-based baseline, suggesting the latent reasoning stage substantially cuts inference cost versus token-heavy chain-of-thought approaches. This is directly relevant to practitioners building faster multimodal chemical-reasoning or molecular-editing pipelines who need throughput gains without sacrificing localization-grounded interpretability.

arXiv · physics.ins-detBuildable

Scan-Coil Delay Causes Anisotropic Signal Loss in Fast 4D-STEM

Electron microscopes blink faster than their steering coils can keep up, blurring the picture.

In advanced electron microscopy called 4D-STEM, a beam of electrons is scanned across a sample point by point, snapping a tiny diffraction pattern at each spot. As scientists push to faster scan speeds (microseconds per point) to reduce damage to delicate samples like proteins, they discovered the electromagnetic coils that steer the beam can't quite keep up in time, causing the beam to lag and smear the image in one direction. The fix is a clever computational trick: compare shifted sub-frames of the data to measure exactly how much lag occurred, then mathematically undo the smear. This matters because it lets scientists recover crisp, trustworthy images from fast, low-dose scans without buying new hardware, which is especially valuable for imaging fragile biological samples.

Technical view

The authors identify an intra-dwell scan-coil settling delay (tens of microseconds) that becomes non-negligible at microsecond dwell times used in fast pixelated-detector 4D-STEM, producing anisotropic signal smearing along the fast-scan axis. They characterize this via direct probe imaging and sub-frame diffraction analysis, then correct it using phase-correlation-based sub-frame alignment applied post-hoc to existing datasets. The correction recovers signal across spatial frequencies, with the largest benefit at large scan steps typical of low-dose biological 4D-STEM. Because it requires no hardware modification, practitioners can apply this correction retroactively to archived fast-scan datasets to improve reconstruction fidelity.

arXiv · q-bio.GNBuildable

Static analysis-guided agentic AI translation enables Rust as a full stack bioinformatics language

AI rewrote a scientist's clunky old bioinformatics code into Rust — 80x smaller, way faster.

Bioinformatics — the software used to analyze DNA and medical images — is often built on decades-old code in languages like Perl or Fortran that few people still know how to maintain, creating security risks and wasted computing resources. The researchers used an 'agentic' AI system (one that can plan and act autonomously, guided by automated code-checking tools) to translate this legacy code into Rust, a modern, fast, and safety-focused programming language. Applied to their own tool, Bascet, the translated version ended up about 80 times smaller, built ten times faster, and ran key operations three times quicker, while also shedding messy external dependencies. This matters because it offers a practical path to modernize critical scientific software without a costly, error-prone manual rewrite.

Technical view

The authors combine static analysis with agentic AI (an LLM-driven system capable of iterative planning and tool use) to systematically translate legacy bioinformatics codebases into Rust, releasing prompts and supporting software for the pipeline. Applied to their NGS/imaging tool Bascet, the translation achieved ~80x reduction in codebase/binary size, ~10x faster build times, and >3x runtime improvement on key steps, while eliminating Unix-specific dependencies for better portability. This demonstrates a reproducible, static-analysis-guided methodology practitioners could apply to other legacy scientific codebases (e.g., Perl, Fortran) facing similar maintainability and safety concerns. The approach suggests a template for using AI agents as verified code-migration tools rather than purely generative assistants.

arXiv · q-bio.PEConceptual

Stochastic Spatial Metapopulation Modelling of HPAI Control and Poultry Restocking on Jolly Island

A simulated bird-flu outbreak on a fake island tests when it's safe to restock chicken farms.

When bird flu hits poultry farms, officials must decide fast how aggressively to cull animals and how soon depopulated farms can safely restart operations. This study builds a computer simulation of a fictional 'Jolly Island' with different farm types (broiler chickens, organic ducks, and others), modeling how the virus spreads through nearby farms, the environment, animal transport, and distance, then testing control strategies like preventive culling and delayed restocking. The simulation showed that culling farms preemptively — before they're confirmed infected — cut the total outbreak burden by about 17%, and confining birds earlier shrank the epidemic further. This matters because it gives policymakers an evidence-based way to weigh the economic and animal-welfare costs of aggressive intervention against the risk of letting an outbreak spread.

Technical view

The authors construct a stochastic spatial SEIR (susceptible-exposed-infectious-recovered) metapopulation model for a synthetic HPAI outbreak, incorporating local, environmental, movement-mediated, and distance-dependent transmission across farm types (Broiler-2, organic duck, Other), alongside reactive/preventive culling, production-specific confinement, and capacity-based restocking rules. Simulations show geographically concentrated epidemics that vary substantially by production class, with preventive culling reducing mean cumulative burden from 16,362.7 to 13,631.9 infectious-farm-days (16.7% reduction), and earlier confinement further reducing epidemic magnitude. The modular framework (synthetic geography, class-specific transmission parameters, intervention triggers) is designed to be adapted to real regional poultry networks for scenario testing. Practitioners in animal health policy could use this structure to benchmark culling thresholds and restocking capacity constraints before applying to real outbreak data.

arXiv · q-bio.NCConceptual

Activity-dependent epidemic spreading on multiscale brain networks predicts Alzheimer's disease progression

Alzheimer's-causing proteins seem to hitch a ride on brain activity itself as they spread.

Alzheimer's and similar diseases progress as toxic misfolded proteins spread from one brain region to connected ones, almost like an infection traveling along neural highways. Prior math models of this spread ignored the fact that neurons firing electrical signals actually helps push these proteins along, something lab experiments have shown. This paper builds a model that links neuronal activity to protein-spreading dynamics (borrowed from epidemic math used to study diseases like flu), finding a tipping point that determines whether a tiny pathological seed grows into full-blown spread, and how activity redirects which brain areas get hit first. This matters because it could help predict individualized disease progression and identify why certain brain networks are more vulnerable, potentially informing where and how to intervene.

Technical view

The authors couple a generic node-activity process to susceptible-infected-susceptible (SIS) epidemic dynamics on brain connectomes, deriving an epidemic threshold that governs whether pathological protein seeds propagate, and showing that a dominant network eigenmode determines the spatial origin of growth. Analytical approximations quantify how neuronal activity shifts this threshold and reshapes spreading trajectories by mixing structural network modes (eigenvectors of the connectivity matrix). For networks with multiscale (hierarchical/modular) structure, they decompose activity-driven effects into contributions from regional mean activity versus within-region variability, enabling attribution of spreading changes to specific network scales. This provides a mechanistic, mode-decomposition framework that researchers could apply to patient-specific connectomes and activity data (e.g., fMRI, EEG) to predict individualized Alzheimer's progression patterns.

arXiv · q-bio.PEConceptual

Dynamics of fluctuating populations in multi-state switching environments

Bacteria don't just face feast or famine — they navigate many shades of 'okay, not great.'

Microbes living in changing environments — think gut bacteria or lab cultures — often face swings between plenty of food and scarcity, and scientists have long modeled this as a simple on/off 'feast or famine' switch. But real environments are messier, with many in-between levels of resource availability, not just two extremes. This paper studies how two competing bacterial strains, one growing slightly slower than the other, fare when the environment cycles through several intermediate states rather than just two, each state offering a different capacity for how many microbes it can support. This matters because understanding these richer fluctuation patterns could reveal why weaker competitors sometimes survive or even thrive in fluctuating real-world conditions, with implications for microbiome ecology and evolution.

Technical view

The authors extend classic binary switching-environment models (used to study feast-famine population dynamics) to a multi-state stochastic framework where the environment transitions among a finite number of intermediate states, each with its own carrying capacity, better approximating experimentally observed gradual nutrient fluctuations. They analyze competitive dynamics between two strains with slightly different growth rates under this richer switching process, likely deriving conditions (switching rates, state structure) under which the slower strain persists or is driven extinct. This generalizes prior two-state (Moran-type or telegraph-process) population models to a broader class of environmental noise, providing a more realistic null model for microbial competition. Researchers modeling real fluctuating ecosystems (e.g., gut microbiota, chemostats) could apply this framework to test how granularity of environmental variation affects coexistence outcomes.

arXiv · cs.LGBuildable

Task- and dataset-specific information in protein language models

AI models of proteins are usually read from their 'last thought' — but earlier layers know more.

Protein language models are AI systems trained on huge databases of protein sequences, similar to how ChatGPT is trained on text, and they convert amino acid sequences into number-based representations used for tasks like predicting protein structure or function. Everyone conventionally uses the output from the model's very last processing layer, treating it like the model's 'final answer,' but nobody had carefully checked whether that's actually the best layer to use. The researchers tested 13 different protein language models across 15 different tasks, training simple predictor 'probes' on the outputs of each internal layer, and found that the last layer is rarely the most informative one. This matters because it means researchers using these models for drug discovery or protein engineering may be leaving performance on the table by defaulting to the last layer instead of the best one.

Technical view

The authors systematically probe 13 protein language models (PLMs) across 15 downstream tasks drawn from 11 datasets, training linear/shallow probes on embeddings extracted from every intermediate layer rather than only the conventional final layer, and supplement this with latent-space geometry characterizations to estimate embedded information content. The key finding: final-layer embeddings rarely yield the best downstream task performance, challenging the field's default convention. This suggests a practical, low-cost intervention — layer selection via lightweight probing — that practitioners can apply to existing pretrained PLMs to improve downstream performance without retraining, and motivates further mechanistic interpretability work on what different PLM layers encode.

arXiv · stat.MEConceptual

Testing the limits of past-adapted explanations by post-endpoint randomisation: anticipatory EEG as a worked case

A clever statistical trick tests whether brain-wave predictions are cheating by peeking at the future.

When scientists build a model that predicts a brain signal (like an EEG wave that ramps up in anticipation of an event), it's hard to know if the model is genuinely using only past information, or if it's secretly benefiting from information that technically comes later. This paper proposes a rigorous testing method: lock in the prediction endpoint before randomly varying the delay to the actual event, so that if the model truly only used past-available information, its predictions shouldn't be able to reliably track that later-randomized delay. Applying this to a specific anticipatory brain wave (called contingent negative variation), they build in strict safeguards against accidental data leakage to make the test trustworthy. This matters because it offers neuroscientists (and other predictive modelers) a principled way to catch models that look good on paper but are secretly relying on information they shouldn't have access to.

Technical view

The paper introduces 'Level II-A,' a design-based causal inference framework that distinguishes model fit from information sufficiency by using post-endpoint randomization: a pre-event endpoint is committed before the delay-to-imperative-event is randomized, turning the later-assigned delay into a negative-control probe. Under a 'past-adapted factorization,' any model using only pre-commitment information should be unable to systematically order the endpoint by the randomized delay; violation of this indicates the model leveraged information beyond what was legitimately available. Applied to anticipatory EEG (contingent negative variation), the framework combines leakage-safe preprocessing with a frozen, label-blind comparator and retained-sample qualifications to produce a confirmatory residual test. Researchers building predictive models from time-series neural data can adopt this design to rigorously test sufficiency claims rather than relying on fit-based validation alone, which is vulnerable to inadvertent information leakage.

arXiv · q-bio.NCConceptual

The Rosetta Stone and Levels of Principled Inference to the Experience of Another Mind

Could a math formula ever let you truly know what it's like to be someone else's mind?

Philosophers have long wrestled with the 'problem of other minds': you know your own inner experience directly, but you can only ever infer someone else's from the outside, leaving what the authors call an 'acquaintance gap.' Some researchers hope that by describing conscious experience mathematically — essentially finding a universal translator or 'Rosetta Stone' for subjective experience — we could bridge that gap. This chapter examines two competing mathematical frameworks for consciousness, the Qualia Structure Paradigm and Integrated Information Theory, asking what each would actually let us conclude about another being's inner experience if it succeeded on its own terms. This matters for debates about animal consciousness, AI sentience, and medical assessments of unresponsive patients, where knowing whether 'someone is home' has real ethical stakes.

Technical view

The chapter conducts a comparative philosophical analysis of two structuralist theories of consciousness — the Qualia Structure Paradigm (Qstr) and Integrated Information Theory (IIT) — evaluating their capacity to license principled inference about another system's phenomenal experience via a mathematical 'Rosetta Stone' translating experiential content into formal structure. Qstr is characterized as proceeding inter-phenomenally, aiming to exhaustively characterize experience through its internal relational structure, while IIT is presumably contrasted on its integrated-information-theoretic mechanism (the abstract cuts off before full elaboration). The analysis likely interrogates whether structural isomorphism between a formal model and a system's causal/relational architecture is sufficient to warrant inference to genuine phenomenal states, bearing on consciousness-detection frameworks proposed for AI systems, animals, or disorders of consciousness. Readers building or evaluating consciousness-measurement frameworks (e.g., IIT's Φ metric) would find this a rigorous epistemic critique of what such measures can and cannot license inferentially.

bioRxiv · molecular biologyConceptual

Z-AAT impairs organelle homeostasis and reduces adaptive response to lipids in alpha-1 antitrypsin deficiency models

A misfolded liver protein doesn't just clump — it quietly wrecks cells' power plants and fat processing.

Alpha-1 antitrypsin deficiency is a genetic disease where a faulty version of a liver protein, called Z-AAT, misfolds and gets stuck inside liver cells instead of being released into the blood. Researchers grew both flat liver cell cultures and tiny 3D lab-grown mini-livers (organoids) from patients' own cells to watch what the stuck protein does over time. They found it doesn't just pile up — it also causes fat to accumulate, damages mitochondria (the parts of the cell that make energy), and forces cells to rely more on sugar than fat for fuel. This matters because it reframes the disease as a broader metabolic breakdown, not just a protein traffic jam, opening new angles for treatment.

Technical view

Using Z-HepG2 cells and ZZ patient-derived hepatic organoids, the authors combined transcriptomic and proteomic profiling with functional assays of mitochondrial and peroxisomal dynamics to characterize downstream effects of Z-AAT polymer accumulation. Z-AAT expression reduced protein secretion and drove lipid accumulation alongside mitochondrial structural abnormalities, an increase in mitochondrial number, and impaired respiratory capacity, with metabolic profiling showing reduced oxidative phosphorylation and a compensatory shift toward glycolysis. The organoid model provides a patient-relevant platform for testing whether restoring mitochondrial or lipid handling mitigates AATD-associated liver disease. This positions organelle-level metabolic dysfunction, not just proteotoxic stress, as a therapeutic target.

bioRxiv · cell biologyConceptual

Isotype specific loss of HP1α but not of HP1β uncovers genomic regions that behave as HP1α-dependent common fragile sites

Losing one copy of a genome-guarding protein leaves specific chromosome spots prone to snapping.

Cells have a family of three related 'HP1' proteins that help pack DNA safely and keep chromosomes stable during division. This study knocked out each HP1 type one at a time in cells and then stressed their DNA-copying machinery with a chemical, checking under a microscope for broken chromosomes. They discovered that losing just one specific version, HP1-alpha, causes breaks at particular genome locations, while losing a very similar cousin, HP1-beta, does not — showing these lookalike proteins actually have distinct, non-swappable jobs. The mechanism seems to be that HP1-alpha loss slows down the process of copying DNA, making certain fragile regions more likely to snap under stress, which is relevant to understanding cancer-related chromosome instability.

Technical view

The authors performed isotype-specific inactivation of HP1α versus HP1β across multiple cell lines and quantified chromosomal breakage on metaphase spreads with and without aphidicolin-induced replication stress. HP1α loss, but not HP1β loss, significantly increased breaks on chromosome arms and within pericentromeric heterochromatin, and was mechanistically linked to reduced replication fork velocity. This defines a set of genomic loci that behave as HP1α-dependent common fragile sites, distinct from classical fragile sites, giving researchers a new isotype-specific readout for probing heterochromatin's role in replication stress and genome instability.

bioRxiv · cell biologyBuildable

Streamlining large-scale high-resolution electron tomography with VolWeaver

A new imaging trick zooms into single molecules inside a parasite without losing the big picture.

Scientists studying cells under powerful microscopes face a tradeoff: some methods show incredible molecular detail but only in a tiny patch of tissue, while others capture a whole cell but blurrily. This project built a new workflow, called VolWeaver, that lets researchers pick a specific tiny region inside a malaria parasite and image it at near-atomic sharpness while still keeping several micrometers of surrounding cellular context visible. It works by carefully optimizing a technique called transmission electron tomography, which takes many angled snapshots of a sample and reconstructs them into a 3D volume. This matters because it finally lets scientists connect fine molecular structures to the broader architecture of the cell they live in, which is crucial for understanding how parasites like the one causing malaria actually work.

Technical view

VolWeaver is an optimized transmission electron tomography (TEM) pipeline for resin-embedded, large-scale volume imaging that targets specific regions of interest at nanometer resolution while retaining several micrometers of surrounding cellular context, applied here to malaria parasites. It bridges the resolution/field-of-view gap between single-particle cryo-EM/cryo-ET and lower-resolution volume EM (e.g., resin-embedded SEM), enabling correlated multiscale analysis within one specimen. Practitioners working with resin-embedded pathogens or organelle-scale structures could adopt the workflow to link molecular-scale tomographic detail directly to whole-cell ultrastructure without switching modalities or samples.

bioRxiv · developmental biologyConceptual

Polarized F-actin establishes cell interactions required for the formation of a stem cell niche

Fruit-fly testis cells use a directional actin 'push' to build the nursery where stem cells live.

Stem cells throughout the body depend on a specialized support structure called a niche, but scientists rarely get to watch a niche actually form because it's usually deep inside developing tissue. Using the fruit fly testis as a visible model system, researchers tracked a cell scaffolding protein called F-actin, which becomes lopsidedly concentrated at specific cell-to-cell contact points as the niche assembles. To test whether this lopsided pattern actively drives cell movement or is just a byproduct of cells sticking together, they used a light-controlled ('optogenetic') tool to switch off a regulator called Rho1 at precise moments and locations, disrupting the actin pattern on demand. Doing so broke normal niche formation, showing that the directional actin arrangement is a cause, not just a consequence, of how stem cell niches get built.

Technical view

Using live imaging of Drosophila testis niche formation, the authors show that F-actin polarizes to specific cell-cell interfaces during niche assembly, and they test causality using optogenetic disruption of Rho1 to acutely and locally perturb cortical F-actin with tissue and temporal precision. Rho1-mediated loss of F-actin polarization produced defects in niche formation, indicating that polarized actin actively directs niche cell positioning/motility rather than merely reflecting adhesion-driven sorting. This establishes an in vivo, optogenetically tractable system for dissecting the cytoskeletal mechanics of stem cell niche morphogenesis, applicable to studying niche construction principles more broadly.

bioRxiv · biochemistryConceptual

Conformational flexibility of soybean lipoxygenase is coupled to crystal solvent content in serial crystallography

An enzyme crystal's own water content secretly dictates how much the protein inside can wiggle.

X-ray crystallography lets scientists see the 3D shape of proteins, but it usually freezes them into one rigid pose, when in reality proteins constantly flex and move. This study used an extremely fast, high-powered X-ray laser technique to look at many tiny crystals of a plant enzyme called lipoxygenase, hoping to catch its natural range of motion. While analyzing the data, they noticed the crystals weren't all identical — some had more water packed between the protein molecules than others — and this water content directly changed how flexible the protein appeared to be. By carefully sorting the data into two groups based on this hidden variation, they solved two distinct structures from a single experiment, revealing that the crystal's surroundings, not just the protein itself, shape what scientists see.

Technical view

Using serial femtosecond crystallography (SFX) at the LCLS free-electron laser on soybean lipoxygenase-1 microcrystal slurries, the authors identified unit-cell polymorphism arising from both indexing ambiguity (pseudo-tetragonal lattice symmetry) and genuine non-isomorphism tied to differing crystal solvent content. By combining unit-cell clustering with systematic reindexing, they deconvolved the mixed dataset into two polymorphs and solved two independent structures showing distinct conformational states from a single experiment. This provides a general data-processing strategy—clustering plus reindexing—for extracting hidden conformational heterogeneity from SFX datasets that would otherwise be merged and averaged away, relevant to anyone doing room-temperature/physiological serial crystallography.

bioRxiv · biochemistryConceptual

Location-dependent proteomics of the aorta reveal an atherosclerotic disease gradient shaped by hemodynamics

Where blood swirls roughly in an artery is exactly where its proteins start looking diseased.

Fatty plaques that clog arteries don't form randomly — they tend to show up at spots where blood flow gets turbulent or disturbed rather than flowing smoothly. Researchers wanted to know what's different, at the protein level, between plaque-prone and plaque-resistant spots in the same blood vessel, but this required measuring proteins in extremely small tissue samples, which used to be technically impossible. Using improved, highly sensitive lab equipment (mass spectrometry) that can now analyze tiny amounts of tissue, they compared proteins from plaque-forming branch points versus healthy-looking curves of the aorta in mice bred to develop atherosclerosis. This location-by-location protein mapping reveals a gradient of disease-related changes tied directly to how blood physically moves through the vessel, which could point to new early-warning markers or treatment targets.

Technical view

The study performed site-resolved proteomics on aortic arch tissue from Western-diet-fed ApoE-/- mice, comparing plaque-prone regions at major branch points and the inner curvature (disturbed flow) against visibly healthy, flow-protected regions, using LC-MS/MS enabled by recent advances in low-input mass spectrometry. This generates a spatial proteomic map linking local hemodynamic conditions to site-specific protein signatures of atherosclerotic disease progression within a single animal. Researchers can use this dataset to identify candidate mechanotransduction pathways or biomarkers specific to flow-driven plaque initiation, and the low-input MS approach itself is reusable for other small-tissue-sample proteomic studies.

bioRxiv · bioengineeringBuildable

NIR-II squeezed light-field microscopy enables high-speed volumetric imaging of deep-tissue dynamics in vivo

A squeezed-light camera trick films a beating heart in 3D, 600 times a second, through living tissue.

Seeing fast 3D movement deep inside living tissue is hard because normal volumetric microscopes either have to scan slowly point-by-point, or they split their camera's limited pixels across many viewpoints at once, losing detail. This is especially tough in a useful infrared wavelength band (NIR-II) that penetrates tissue well but whose cameras are small and noisy. The researchers built a new microscope, NIR-II squeezed light-field microscopy, that optically twists and compresses multiple viewing angles together before they hit the camera, so the limited pixels get used far more efficiently while still capturing enough information to reconstruct a 3D image. The result is volumetric video fast enough — up to 600 3D frames per second — to capture things like a beating heart in real time, without needing any fluorescent labels or dyes.

Technical view

NIR-II SLIM introduces an optical rotate-and-compress step for multi-perspective light-field views prior to detection, maximizing effective use of a small-format, high-noise InGaAs sensor in the second near-infrared window while preserving the spatial information needed for 3D reconstruction. The system achieves volumetric acquisition rates up to 600 volumes/s at a 512×512 reconstructed lateral sampling grid, demonstrated on label-free, four-dimensional imaging of cardiac dynamics in vivo. This addresses the core pixel-budget bottleneck of snapshot light-field imaging in NIR-II, offering a hardware-optical (rather than purely computational) route to high-speed deep-tissue volumetric imaging that could be adapted to other fast physiological dynamics.

bioRxiv · bioinformaticsRunnable

Detection of Spatially Aberrant Cells in Spatial Transcriptomics Data by Conformal Prediction

An algorithm spots tissue cells that broke their normal honeycomb formation — a red flag for disease.

Healthy tissues like skin or gut linings have cells arranged in tidy, repeating hexagonal patterns, and when disease starts, that neat organization breaks down into oddly placed cells with abnormal gene activity. The researchers built a computer tool called SPADE that combines two types of data — regular single-cell gene readouts and spatial maps showing where cells actually sit in tissue — to learn what 'normal' spatial-and-genetic patterns look like. It then flags locations that don't fit that learned normal pattern, while also reporting how confident it is in each flag, so users know which anomalies to trust versus which might be noise. This gives researchers a statistically principled way to automatically catch early signs of disease-related tissue disorganization instead of relying on visual inspection.

Technical view

SPADE integrates scRNA-seq and spatial transcriptomics via a variational autoencoder combined with Gaussian mixture modeling to jointly learn cell-type embeddings and perform spatial deconvolution, then applies conformal prediction to yield uncertainty-calibrated detection of spatially aberrant spots. The conformal prediction layer provides formal statistical coverage guarantees on aberrancy calls rather than ad hoc thresholding, and the method reportedly outperforms existing approaches in benchmark validation. As a computational framework it's directly usable by anyone with paired scRNA-seq/spatial transcriptomics datasets to screen for disease-associated spatial disorganization, likely released as an installable analysis pipeline.

bioRxiv · cancer biologyConceptual

Replication fork plasticity is a therapeutic vulnerability in acute myeloid leukemia

Leukemia cells' DNA-copying machinery can be sabotaged to make cancer drugs work better.

Acute myeloid leukemia (AML) is caused by bone marrow cells multiplying out of control, and current treatments are often harsh and don't always work. This research looks at drugs called PARP inhibitors, which block a DNA-repair protein, and finds they only work well in some genetic subtypes of AML. Using detailed imaging of individual DNA strands as cells copy their genome, the researchers show that in drug-sensitive leukemia cells, the treatment first speeds up DNA copying dangerously fast, then causes the DNA strand to snap later in the same cell cycle. Understanding this timeline could help doctors predict which patients will actually benefit from these drugs.

Technical view

The study combines single-cell and single-molecule DNA fiber assays with damage-signaling readouts to dissect replication fork dynamics under PARP inhibition and cytarabine in AML models stratified by fusion genotype. In PARPi-sensitive backgrounds (RUNX1-RUNX1T1, PML-RARα), PARP inhibition deregulates RECQ1-dependent fork restart, producing an initial phase of fork acceleration followed by fork breakage within the same S phase, whereas KMT2A-rearranged (resistant) cells apparently evade this cascade. This mechanistic ordering — accelerated restart preceding catastrophic breakage — offers a candidate biomarker/timing window for combination scheduling and could guide biomarker-driven PARPi trial stratification in AML.

bioRxiv · cancer biologyConceptual

p140Cap enhances breast cancer chemosensitivity by limiting an ABCC1-enriched stem-like compartment via β-Catenin inhibition

One protein makes breast cancer cells drink in more chemo instead of pumping it back out.

Breast cancer chemotherapy doesn't work equally well for everyone, partly because some tumor cells act like resistant 'stem cells' that survive treatment and cause relapse. This study looks at a protein called p140Cap and finds that it helps chemotherapy drugs like doxorubicin stay inside cancer cells longer, causing more DNA damage and killing more cells. It works by blocking a signaling pathway (Wnt/β-Catenin) that would otherwise boost a pump protein (ABCC1) that stem-like cells use to flush the drug back out. This suggests that boosting p140Cap activity, or blocking the Wnt pathway it controls, could make existing chemotherapies more effective against hard-to-treat breast cancers.

Technical view

Using preclinical and patient-derived HER2-positive and triple-negative breast cancer models, the authors show p140Cap increases intracellular doxorubicin retention and downstream DNA damage/apoptosis by restraining a doxorubicin-negative, ABCC1-high side population enriched for stem-like cells. Mechanistically, p140Cap inhibits β-Catenin signaling, which otherwise drives ABCC1 (a drug-efflux transporter) expression; constitutively active β-Catenin reverses the chemosensitizing phenotype, while pharmacological Wnt/β-Catenin inhibition phenocopies p140Cap's effect. This positions the p140Cap–β-Catenin–ABCC1 axis as a druggable node for combination strategies aimed at depleting the chemoresistant stem-like compartment.

bioRxiv · plant biologyConceptual

Concomitant post-translational repression of Arabidopsis PIP1 aquaporins upon the loss of major PIP2 isoforms

Knock out one type of plant water channel, and a related one quietly gets destroyed too.

Plant cells regulate water flow using channel proteins called aquaporins, which come in two related families, PIP1 and PIP2, sitting in the cell membrane. Researchers genetically removed several PIP2 versions in the mustard-family plant Arabidopsis and unexpectedly found that PIP1 protein levels also crashed, even though the genetic instructions (mRNA) for making PIP1 were untouched. This means the loss happens after the protein is made, not because the plant stopped producing it, hinting that PIP1 proteins need PIP2 partners to avoid being tagged for destruction by the cell's protein-disposal systems. It's a reminder that in biology, removing one gene can ripple through and silently take out a partner protein via a completely different mechanism.

Technical view

In a pip2;1 pip2;2 pip2;4 pip2;6 pip2;7 quintuple mutant, Arabidopsis PIP1 protein abundance is strongly reduced despite unchanged PIP1 steady-state transcript levels and polysome loading, indicating a post-translational (not transcriptional/translational) mechanism. Intermediate mutant combinations (pip2;1 pip2;2 and pip2;1 pip2;2 pip2;7) show graded residual PIP1 protein (60% and 20%), consistent with dose-dependent stabilization of PIP1 by PIP2 heteromerization. The authors are positioned to test which degradation pathway — ER-associated ubiquitin-proteasome degradation or an alternative route — clears unpartnered PIP1, informing models of aquaporin heteromer-dependent quality control at the plasma membrane.

bioRxiv · plant biologyBuildable

Field Modeling Study of Yield Response and Nitrate Leaching with Manure Application and Deficit Irrigation for Maize-Fallow-Wheat Rotation

Cow manure plus a bit less irrigation water can grow the same wheat with less nitrate pollution.

Farmers in dry regions face a tricky balance: use enough fertilizer and water to grow good crops, without wasting water or letting nitrogen runoff pollute groundwater. This two-year field study in Pakistan compared spreading dairy manure versus synthetic urea fertilizer, combined with either full or reduced ('deficit') irrigation, across a wheat-fallow-maize rotation. They buried special sensors deep in the soil to directly measure how much nitrate leaked downward past the root zone, and used a computer model to estimate daily water loss. The combination of manure with reduced irrigation turned out to be a promising sweet spot, suggesting practical changes farmers could make to protect water quality without sacrificing yield.

Technical view

The two-year field trial crossed dairy manure (50 Mg/ha annually, N-equivalent to recommended urea) against sole urea fertilization with 100% vs 75% ETc irrigation in a wheat-fallow-maize rotation, using suction lysimeters at 1.2 m depth for direct nitrate-N leachate measurement and HYDRUS-1D modeling for daily deep percolation. A significant manure × irrigation interaction affected yield, nitrate-N leaching, and soil quality, with manure under deficit irrigation improving wheat outcomes — data practitioners could use to parameterize regional HYDRUS-1D models or design manure/deficit-irrigation protocols for semi-arid cereal systems.

bioRxiv · plant biologyRunnable

Plant functional defects experienced upon growth under Per-/Poly- fluoroalkyl substances (PFAS) conditions

'Forever chemicals' in water can stunt bean sprouts before they even get their first leaf.

PFAS, nicknamed 'forever chemicals' because they don't break down in the environment, are showing up in soil and water everywhere, but scientists don't fully know how they affect the crops we eat. This study grew mung beans in water spiked with two common PFAS chemicals (PFOA and PFOS) at different concentrations, to see how the plants responded. At very high doses, PFOA badly delayed seed sprouting, leaf growth, and root hair formation, while a wider range of doses caused temporary stunted growth that plants partly recovered from over time. At the highest tested dose, plants ended up smaller and lighter, showing that PFAS contamination can measurably harm crop development, and worse at higher concentrations.

Technical view

Hydroponic mung bean (Vigna radiata) seedlings were exposed to PFOA and PFOS across a dose range (5-500 µM, plus a 1 mM PFOA high-dose condition), with developmental endpoints including germination timing, leaf emergence, root hair formation, wet weight, and leaf area/biomass tracked over multiple timepoints (48 h onward). PFOA at 1 mM produced severe, more pronounced impairment than PFOS at equivalent exposure, while the 5-500 µM range induced transient growth stunting with partial recovery, and 500 µM significantly reduced biomass without affecting an unspecified additional trait. The dose- and compound-specific phenotyping establishes baseline toxicity thresholds usable for follow-up mechanistic work (e.g., uptake/transport assays) or risk assessment in legume crops.

bioRxiv · microbiologyConceptual

A periplasmic regulator establishes adaptive impermeability to control carbapenem entry

Bacteria have a fast pre-emptive shield against antibiotics, before they even bother changing their genes.

Antibiotic resistance in bacteria like Pseudomonas aeruginosa often depends on how easily drugs can get through the bacterial outer wall, through pore proteins called porins. Scientists discovered a small protein, PtrA, that makes bacteria resistant to a key antibiotic (imipenem, a carbapenem) not by getting rid of the pore, but by physically interfering with it to block drug entry. This response is triggered by metal signals (zinc and copper) and kicks in fast, acting as an early defense before the bacteria's slower strategy of shutting down pore production even begins. This reveals a previously hidden, quick-response layer of antibiotic resistance that could be a new target for drugs designed to keep bacteria vulnerable.

Technical view

PtrA is a small periplasmic protein identified as a regulator of OprD-dependent carbapenem permeability in Pseudomonas aeruginosa, conferring imipenem resistance without reducing OprD porin abundance, instead associating with OprD-containing membrane complexes to physically restrict drug entry. This mechanism is mechanistically distinct from and precedes the transcriptional CzcRS-mediated repression of oprD, and is triggered by zinc/copper as physiological signals, defining a two-phase adaptive resistance model: rapid post-translational periplasmic gating followed by slower transcriptional porin downregulation. PtrA represents a novel, non-transcriptional resistance node that could be targeted to resensitize P. aeruginosa to carbapenems, independent of existing porin-expression-focused resistance mechanisms.

bioRxiv · microbiologyBuildable

Environmental and spatiotemporal drivers of marine microbial communities from Antarctic and Subantarctic water masses

Southern Ocean microbes form distinct neighborhoods, and where you look changes what you find.

The Southern Ocean around Antarctica is full of microscopic marine life that forms the base of the food web, but most studies of it have focused on easy-to-reach coastal spots. This research instead sampled open ocean and sub-Antarctic waters, sequencing DNA (specifically a gene called 16S rDNA used to identify bacteria) from four different locations, and matched it with ocean measurements like temperature and currents. They found that microbial communities differed a lot between regions and even between nearby localities, shaped by both local water conditions and the physical difficulty of organisms spreading between distant sites. A few specific microbial groups showed up consistently everywhere, suggesting they are especially well-adapted to the harsh, varied conditions across this remote ocean region.

Technical view

The study used 16S rDNA high-throughput sequencing paired with oceanographic data to characterize marine microbial community composition across four localities spanning two Subantarctic sites (Magellan Strait, Beagle Channel) and two Antarctic open-sea regions (Eastern Indian, Central South Pacific). Results show significant inter-regional and inter-locality divergence in taxonomic composition, alpha diversity, and enriched taxa, attributable to both local environmental filtering and dispersal limitation, while taxa including Clade Ia, Amylibacter, and NS5/NS2b marine groups showed broader distribution across sites. The dataset extends Southern Ocean microbial biogeography beyond coastal-focused sampling and provides a baseline for modeling dispersal-vs-environment drivers of community assembly in under-sampled circumpolar and subantarctic waters.

bioRxiv · molecular biologyBuildable

A multi-omics view of wax synthesis in the wild cochineal bug, Dactylopius opuntiae

Scientists sequenced the cactus bug that makes red dye, to find the genes behind its waxy coat.

The cochineal bug is a small insect known for its dense white waxy covering and, in a related species, for producing red dye, but it's also a pest that damages prickly pear cacti. This study built a full genetic blueprint (genome) of the wild cochineal bug for the first time, alongside data on which genes are active and which proteins are made. The researchers specifically hunted for a family of genes called fatty acyl reductases (FARs), which are known to help insects build their protective wax coatings, and found 26 of them, some of which appear to have multiplied and specialized within this species. This genetic map gives scientists the tools needed to understand — and potentially disrupt — how this pest builds its wax armor, which could inform pest control strategies for cactus farming.

Technical view

The authors report a 359-Mb de novo genome assembly for Dactylopius opuntiae alongside transcriptomic and proteomic data, identifying 26 fatty acyl reductase (FAR) genes central to epicuticular wax biosynthesis. Phylogenetic reconstruction across Coccoidea (scale insects) reveals lineage-specific tandem expansions of FAR genes in D. opuntiae, implicating gene duplication as a driver of this species' distinctive dense wax phenotype. This multi-omics resource enables comparative genomic studies of wax biosynthesis across scale insects and provides candidate gene targets for pest-control strategies targeting the insect's protective wax layer.

bioRxiv · molecular biologyBuildable

A Robust and Scalable Workflow for the Production of Circular Single-Stranded DNA for Genome Engineering Applications

A cheap, reliable recipe for growing circular DNA rings used to edit genomes.

Scientists often need long loops of single-stranded DNA (a floppy, one-sided version of the usual double helix) as tools for editing genomes, building tiny DNA structures, or making diagnostic tests. These circular loops are tougher than straight DNA strands because cells' natural DNA-chewing enzymes have a harder time degrading a ring than a strand with loose ends, and they can be made much longer than DNA can be chemically synthesized. Until now, making them required expensive specialty kits or ordering custom DNA from a company. This paper lays out a step-by-step lab recipe using a virus-derived tool called a phagemid (basically hijacking a bacterial virus's DNA-copying machinery) to mass-produce pure circular DNA cheaply, making a widely useful material accessible to any lab.

Technical view

The authors present a scalable protocol for producing circular single-stranded DNA (cssDNA) using an M13 phagemid system, addressing the field's dependence on costly commercial synthesis or specialized reagents for genome-editing and nanotechnology applications. cssDNA's circularity confers exonuclease resistance and enables generation of long, sequence-defined constructs beyond chemical synthesis limits. The workflow appears optimized for reproducibility and yield at bench scale, likely detailing phage propagation, ssDNA extraction, and purification steps. Researchers building genome-editing templates (e.g., for HDR or prime editing donors) or DNA nanostructures could adopt this protocol directly to replace outsourced synthesis.

bioRxiv · molecular biologyBuildable

Genome-wide mapping of helicase-generated ssDNA reveals Hrq1 activity at RNA polymerase III-transcribed genes

Researchers built a genome-wide GPS to catch DNA-unwinding enzymes in the act.

DNA helicases are molecular motors that unzip the double helix so it can be copied, repaired, or read — but finding exactly where in the genome they're working has been hard because their action leaves no visible trace. The researchers solved this by fusing a helicase to an enzyme that chemically tags any exposed single-stranded DNA, leaving tiny 'footprints' (mutations) wherever the helicase had unzipped DNA, which they then read out by sequencing the whole genome. Using this trick on a yeast helicase called Hrq1 (related to a human gene linked to disease), they discovered it works heavily at genes read by a specific gene-copying machine, especially the small genes that make transfer RNA. This gives scientists a general new method to map where any helicase acts across an entire genome at near-single-letter precision.

Technical view

The authors fuse DNA helicases to activation-induced cytidine deaminase (AID), which deaminates cytosines only in transiently exposed ssDNA, converting helicase unwinding events into strand-specific mutational footprints readable by whole-genome sequencing at near-nucleotide resolution. Applying this to S. cerevisiae Hrq1 (a RecQ4-family helicase and functional homolog of human RECQL4, implicated in Rothmund-Thomson syndrome) produced the first genome-wide activity map, revealing strong enrichment at RNA Pol III-transcribed loci, particularly tRNA genes, with strand bias favoring the non-template/transcript strand. This AID-fusion approach is generalizable to other helicases and translocases, offering a template for in vivo mapping of ssDNA-generating enzymes beyond ChIP-based methods.

bioRxiv · molecular biologyBuildable

Studying the effect of conserved tyrosine phosphorylation within SH2 domains

A chemical tag on a signaling protein can silently rewire which partners it grabs.

Cells constantly send messages by attaching phosphate tags to proteins, and one common 'reader' of these tags is a protein module called an SH2 domain, which itself can also get tagged. This study asks what happens when the reader gets tagged too — does it change what the reader can read? Using a screening method (a modified dot blot, essentially a way to test many protein interactions on a membrane at once) with mutations that mimic permanent tagging, the researchers found that tagging one specific spot on a protein called PTPN11 changes its pickiness, making it bind fewer of its usual partners. This matters because PTPN11 and similar proteins are central hubs in cell growth signaling, and disrupting them is linked to cancer and developmental disorders, so understanding this extra layer of control could reveal new ways to intervene.

Technical view

SH2 domains recognize phosphotyrosine motifs to mediate signaling protein-protein interactions, but the regulatory effect of phosphorylation occurring within the SH2 domain itself (as opposed to on its ligand) was poorly characterized. The authors used phosphomimetic mutagenesis combined with a modified dot blot binding assay to interrogate two conserved intra-domain tyrosine phosphorylation sites across PTPN11-N, LYN, and SYK-C SH2 domains. They found that phosphomimicking Y63 in the PTPN11 N-terminal SH2 domain alters binding specificity, reducing affinity for physiologically relevant substrates — implicating this residue in an autoregulatory layer distinct from canonical ligand-based control. This establishes a scalable assay for probing SH2 phosphoregulation that could be extended to other domains implicated in RASopathies and cancer signaling.

bioRxiv · cell biologyBuildable

Development and validation of an SDA-500 Anopheles stephensi cell line for molecular studies

A new lab-grown cell line gives malaria researchers a working model of a key mosquito.

Anopheles stephensi is a mosquito species spreading malaria in cities, but scientists have had almost no lab tools — like cultured cells — to study its biology or engineer genetic controls against it, unlike better-studied mosquito species. This paper reports growing a new, self-sustaining line of cells taken from this mosquito's embryos, then verifying it's really from this species and figuring out its chromosomes. The team also tested different methods for getting foreign genetic material into these cells and found one reagent worked notably better than the common alternative, then used a light-producing reporter system to check which genetic 'on' switches work best in these cells. This toolkit fills a major gap, letting researchers now run genetic experiments directly in cells from this important disease-carrying mosquito instead of guessing from related species.

Technical view

The authors established and validated SDA-500, a novel embryo-derived cell line from Anopheles stephensi, an urban malaria vector for which molecular tools have been lacking. Species and karyotype were confirmed via mitochondrial COI barcoding and cytogenetics (revealing a diploid complement including a Y chromosome). They benchmarked transfection reagents, finding TransIT-PRO outperforms Lipofectamine-based methods, and used dual-luciferase reporter assays to evaluate promoter activity across candidates. This cell line provides a practical in vitro platform for functional genomics and testing genetic control constructs (e.g., gene drives) in A. stephensi, directly usable by vector biology labs for transfection-based screens.

bioRxiv · cell biologyConceptual

The centromere localization domain of kinetoplastid kinetochore protein KKT2 recognizes the free N-terminus of histone H3

An ancient parasite's chromosome-grabbing protein reads a histone's bare tip, not its usual mark.

When cells divide, a molecular machine called the kinetochore grabs each chromosome and pulls it to the right place — normally this machine locks onto a special marker protein (CENP-A) built into the chromosome's DNA-packaging spool. But kinetoplastids, a group of ancient single-celled parasites, lack that marker entirely, so how their kinetochore finds the right spot has been a mystery. This study shows that one of their kinetochore proteins, KKT2, instead directly grabs the very tip of an ordinary packaging protein (histone H3) — and strikingly, even the smallest possible chemical modification to that tip completely blocks the grip. Using precise molecular measurement techniques, the researchers pin down exactly how this alternate recognition works, revealing a completely different solution evolution found for the same essential problem of finding chromosome attachment points.

Technical view

Kinetoplastids lack CENP-A, the centromere-specifying histone H3 variant universal to canonical eukaryotic kinetochores, raising the question of how their divergent kinetochore proteins achieve centromere-specific localization. The authors show via NMR spectroscopy and isothermal titration calorimetry that the centromere localization (CL) domain of Trypanosoma brucei KKT2 structurally resembles a ZZ domain and binds the free N-terminus of histone H3 directly, using an invariant aspartate conserved in known H3-binding ZZ domains. Critically, even N-terminal mono-methylation of H3 abolishes binding, indicating strict recognition of the unmodified free amino group — a mechanism orthogonal to canonical CENP-A-based recruitment. This suggests kinetoplastids evolved an alternative, modification-sensitive histone-reading strategy for centromere specification, offering a comparative model for kinetochore assembly diversity.

bioRxiv · cell biologyBuildable

Optogenetic Control of cAMP Levels and HCN Channels: Implications in Cardiac Physiology and Parkinson's Disease

Flipping a light switch inside cells raised a signaling molecule to fix heartbeats and Parkinson's tremors in mice.

Cells use a molecule called cAMP as an internal messenger to control things like heart rate and certain channels involved in Parkinson's disease. Instead of using drugs, which are blunt and slow, these researchers used optogenetics — inserting a light-sensitive enzyme into cells so that shining light on them precisely raises cAMP levels on demand. They showed that this light trigger sped up the beating of heart cells and, when aimed at a brain region involved in movement in mice with a Parkinson's-like condition, partly restored normal motor function and reduced abnormal spinning behavior. This proves that precisely controlling a single signaling molecule with light, rather than a drug hitting many targets, can meaningfully influence both heart rhythm and Parkinson's-like symptoms, pointing toward more targeted future therapies.

Technical view

The authors express a photoactivated adenylyl cyclase (PAC(S27A)) to optogenetically raise intracellular cAMP with light, then examine downstream effects on HCN channels, which are cAMP-gated and implicated in both cardiac pacemaking and PD pathophysiology. Light-induced cAMP elevation activated HCN4 to increase cardiomyocyte beating rate, while unilateral PAC(S27A) expression in the substantia nigra pars compacta of mice produced light-dependent rotational behavior attenuable by HCN inhibitors, and partially rescued motor deficits in an MPTP-induced PD model with concurrent HCN2 changes. This establishes PAC(S27A) as a tractable optogenetic tool for dissecting cAMP-HCN signaling causally in both cardiac and dopaminergic circuits, with potential application toward HCN-targeted PD or arrhythmia interventions.

bioRxiv · developmental biologyConceptual

In vitro fertilisation and vitrification disrupt embryo mitochondrial function and redox balance that persists into adulthood in mice

IVF and embryo freezing leave a mitochondrial scar on the heart that lasts into adulthood.

IVF (in vitro fertilization) has helped create over 10 million babies, and children conceived this way sometimes show subtle heart differences like altered heart structure and higher blood pressure, but nobody knew why. This study looked at mitochondria — the energy-generating parts of cells — in mouse embryos made by IVF, including ones that were frozen and thawed (vitrified), and tracked what happened as those mice grew into adults. They found that IVF disrupted the embryos' mitochondrial energy processing and their balance of reactive, potentially damaging molecules right from the earliest stage, and traced these disruptions forward to find they persisted in the adult heart. This is the first evidence linking early embryo-stage mitochondrial stress from fertility treatments directly to lasting heart effects in adulthood, suggesting doctors may need to consider long-term cardiovascular monitoring for IVF-conceived individuals.

Technical view

Using the IGS-CD1 mouse model, the authors compared blastocysts derived from natural mating versus IVF (transferred fresh or after vitrification-warming), assessing mitochondrial redox balance and metabolic function at the blastocyst stage and then following offspring hearts into adulthood. IVF reduced total blastocyst mitochondrial content/function and altered redox status, and these early perturbations were traceable to persistent cardiac abnormalities in adult offspring, providing a mechanistic link between preimplantation ART exposure and previously reported ART-associated cardiovascular phenotypes (cardiac remodeling, elevated blood pressure). This is apparently the first study to directly connect blastocyst-stage mitochondrial dysfunction to adult cardiac outcomes in an ART model, offering a framework for testing interventions (e.g., antioxidant supplementation during IVF) to mitigate long-term cardiovascular risk.

bioRxiv · developmental biologyConceptual

Reciprocal mechanochemical feedback couples neural crest migration and neurulation

Migrating cells physically sculpt the very tissue they need in order to migrate — a two-way construction crew.

As a vertebrate embryo's head forms, two things happen almost simultaneously: neural crest cells crawl away from the developing brain tube, and that tube folds itself closed (neurulation). Scientists used to think these were separate, independent programs, but this study shows they're actually in constant physical conversation. As the migrating neural crest cells crawl out, they remodel a scaffold protein called fibronectin in the space between tissues, and this remodeled scaffold both keeps the tissues properly separated and helps the neural tube fold correctly by letting its cells rearrange and squeeze into shape. This remodeling depends on a molecular pair of scissors called MMP14 that only the migrating cells carry, revealing that cell migration and tissue folding aren't separate assembly lines but a feedback loop where each shapes the other — insight relevant to birth defects like neural tube closure failures.

Technical view

The authors demonstrate that cephalic neural crest migration and neural tube closure, long treated as tissue-autonomous processes, are mechanochemically coupled via extracellular matrix remodeling. As neural crest cells delaminate and invade adjacent mesoderm, they remodel fibronectin at the neural crest-neural plate boundary in an MMP14 (membrane-bound metalloproteinase)-dependent manner, generating an ECM interface that both physically separates the tissues and permits radial intercalation and apical constriction in the neural plate necessary for tube closure. Loss of neural crest-specific MMP14 or migration disrupts this ECM remodeling and impairs neurulation, establishing collective cell migration as a mechanical prerequisite for adjacent morphogenesis. This reciprocal feedback model reframes neurulation defects (a major class of human birth defects) as potentially arising from disrupted neural crest-ECM crosstalk rather than neural plate-intrinsic failure alone.

bioRxiv · developmental biologyConceptual

A NON-CANONICAL ROLE FOR NOTCH3 IN BUILDING THE INTESTINAL LYMPHATIC NICHE

A surprise gene helps build the gut's tiny pipes that pump fat into your blood.

Deep inside the lining of your intestine are microscopic lymph vessels called lacteals that soak up fat from digested food, and they only work because tiny muscle cells around them squeeze rhythmically to push the fluid along. Scientists didn't know how the different support cells building this muscle-vessel unit actually talk to each other during development. Using single-cell gene mapping, mouse genetic tracing, and lipid-absorption tests, this team found that a gene called Notch3 acts like a foreman, coordinating separate crews of connective-tissue cells so the muscle wrapping forms correctly. It matters because when this assembly goes wrong, it can cause lymphatic diseases that are currently very hard to treat.

Technical view

The study uses developmental single-cell RNA profiling, Cre-based lineage tracing, and conditional mouse genetics to dissect assembly of the muscular-lacteal complex (MLC) in intestinal villi. Notch3 is identified as a key coordinator between distinct mesenchymal lineages, promoting smooth muscle differentiation within the PDGFRα lineage, even though PDGFRβ lineage cells themselves do not directly contribute to villus smooth muscle — implying a non-cell-autonomous signaling relay. Functional lipid-absorption assays link MLC integrity to physiological lacteal pumping. This establishes a genetic entry point for dissecting mesenchymal crosstalk in lymphatic vessel maturation, relevant to lymphatic dysfunction disorders.

bioRxiv · ecologyConceptual

Sulfur isotopes in hunted ungulates reveal Palaeolithic human mobility patterns in Northern Iberia

Sulfur trapped in ancient hunted bones maps how Stone Age people roamed Spain.

Long before writing, hunter-gatherers in northern Spain moved across the landscape following game and resources, but figuring out exactly how far and how often is hard when all you have is old bones. Researchers measured sulfur, carbon, and nitrogen isotopes — chemical signatures locked into bone collagen that vary depending on the specific soil and region an animal grazed in — from over 900 hunted animal bones spanning nearly 100,000 years. By combining these isotope 'fingerprints' with protein analysis, climate reconstruction, and statistical dating models, they could trace which regions the meat (and therefore the hunters) came from over time. This gives a rare, detailed picture of how prehistoric human mobility patterns shifted with climate and culture, from Neanderthals through early modern humans.

Technical view

The team applied δ34S isotope analysis, cross-validated with δ13C and δ15N and palaeoproteomic species identification, to 905 anthropogenically modified ungulate bone collagen samples from 16 Cantabrian sites spanning MIS 5 to 1 (100–7 ka BP). Sulfur isotopes serve as a geolocation proxy since bedrock-derived sulfate signatures vary spatially, enabling isoscape mapping of animal (and by extension human) provenance. Combined with Bayesian chronological modeling and palaeoclimatic reconstruction, this reconstructs shifts in hunter-gatherer territorial ranges and resource exploitation strategies across the Mousterian-to-Mesolithic transition. The approach offers a replicable isotopic-provenancing framework for other regions with dense faunal assemblages.

bioRxiv · ecologyBuildable

Bridging Ecological Inference and Decision Optimization for Conservation Using Artificial Intelligence

AI that learns population biology AND makes the best call for saving an endangered fish.

When trying to save an endangered species, wildlife managers face a tough tradeoff: models detailed enough to reflect real biology are often too complex to use for picking the best action, while simple decision models ignore important biological nuance. This research fuses two AI approaches — one that builds a rich, data-grounded picture of how a population actually grows and shrinks (an integrated population model), and another that's very good at learning optimal strategies under uncertainty (deep reinforcement learning, the technique behind game-playing AI). Together they let managers get recommendations that are both ecologically realistic and genuinely optimized. They tested this on real conservation efforts to help the endangered Rio Grande silvery minnow, showing it can guide decisions like when and how many fish to release into the wild.

Technical view

The framework couples integrated population models (IPMs), which synthesize multiple data streams (survival, reproduction, abundance surveys) into a unified demographic model, with deep reinforcement learning (DRL) to solve for adaptive management policies over high-dimensional, uncertain state spaces — avoiding the simplification typically required by classical stochastic dynamic programming. Applied to the Rio Grande silvery minnow supplementation program, the IPM-DRL pipeline learns policies directly from the fitted ecological model rather than a reduced-form approximation. This is a template for closing the gap between ecological realism and decision-theoretic optimality in endangered species management, and could generalize to other supplementation or harvest-control problems where population models already exist.

bioRxiv · ecologyRunnable

PAMalytics: a no-code application for structured validation of bioacoustic detections

A no-code app helps scientists double-check what AI 'heard' in wildlife audio recordings.

Researchers now use AI to sift through huge amounts of field audio recordings to detect specific animal calls, but the AI often makes mistakes, so humans still need to listen and confirm each detection before trusting the results. Until now that verification step was done through messy, ad hoc spreadsheets and manual processes prone to error and hard to document. PAMalytics is a free, open-source app that runs right in your web browser (no programming needed) to organize and standardize this human review process. It matters because reliable species monitoring — for conservation, environmental impact assessments, and biodiversity tracking — depends on trustworthy validation, not just a good classifier.

Technical view

PAMalytics is an open-source, local browser-based, no-code application targeting the post-classification validation bottleneck in passive acoustic monitoring (PAM) pipelines. It provides structured workflows for reviewers to confirm or reject automated species-classifier detections, addressing the gap between classifier output and downstream occupancy/detection models that assume validated data. By standardizing and logging the validation process, it improves transparency, reduces transcription/consolidation errors, and creates auditable records — useful for anyone building PAM-based biodiversity monitoring or regulatory reporting pipelines who currently relies on manual spreadsheet review.

bioRxiv · biochemistryBuildable

Functional Differentiation of GH172 Arabinofuranosidases Through Divergent Quaternary Structures

Same enzyme family, three different 3D shapes, three different jobs cutting TB's cell wall sugar.

Mycobacteria — the family that includes tuberculosis — build their tough cell walls using an unusual sugar chain called arabinan, and figuring out how to break it down could help fight these bacteria or study their biology. A gut bacterium was found that can fully dismantle this sugar using a toolkit of enzymes, including three related ones (called GH172) that snip the chain from its ends. This study shows that even though these three enzymes look similar, they cut the chain at different specific points and, surprisingly, assemble themselves into completely different 3D shapes. Using precisely designed test sugars, custom chemical probes, and powerful imaging (X-ray and cryo-electron microscopy), the researchers mapped out exactly how form relates to function — insight useful for designing drugs or biotech tools that target this cell-wall sugar.

Technical view

Three GH172 exo-α-arabinofuranosidases from Dysgonomonas gadei, previously shown to fully degrade mycobacterial arabinan (a component of arabinogalactan and lipoarabinomannan), were characterized against defined synthetic substrates and shown to have distinct linkage specificities. The team developed α-arabinofuranosyl cyclophellitol aziridine-based covalent inhibitors and activity-based probes to interrogate active-site chemistry, and solved X-ray crystal and cryo-EM structures revealing markedly different quaternary assemblies among the three homologues despite shared catalytic fold. This links oligomeric architecture to functional divergence within a single GH family and provides new chemical tool compounds (activity-based probes/inhibitors) usable for studying arabinan-processing enzymes or mycobacterial cell-wall biology more broadly.

bioRxiv · bioengineeringBuildable

Focused framework sampling recovers binding-positive humanized anti-amyloid-β antibodies in a single sorting round

One round of lab sorting turned mouse antibodies into human-ready Alzheimer's drug candidates.

To make an antibody drug from an animal immune response usable in humans, scientists must 'humanize' it — reshape it to look human enough that the immune system won't reject it — while still binding its target tightly, which traditionally takes many slow rounds of trial and error. Here, researchers targeting the amyloid-beta protein linked to Alzheimer's disease combined an efficient antibody-display technique with a smarter, structure-guided way of choosing which parts of the antibody to humanize. They immunized cells, screened billions of antibody variants using fluorescence sorting, and speeply humanized two promising candidates in a single screening round instead of many. This dramatically speeds up an otherwise slow, expensive bottleneck in developing new antibody therapies, here demonstrated for a major Alzheimer's drug target.

Technical view

The workflow combines immune yeast Fab-display (library diversity ~3.5×10^8) with magnetic enrichment and FACS to isolate anti-Aβ1-42 aggregate-reactive clones, then applies structure-guided single-round focused humanization by sampling framework positions predicted to support CDR conformation or VH/VL domain packing, rather than iterative loop-by-loop humanization. Two lead clones, CLAB17 and CLAB45, were converted to full IgG and advanced through this pipeline, achieving binding-positive humanized candidates from a single FACS sorting round. This demonstrates a generalizable, faster alternative to conventional multi-round CDR-grafting/back-mutation humanization for early-stage antibody developability assessment, applicable beyond amyloid-β targets.

bioRxiv · bioengineeringRunnable

Assessing the fractional contributions of static, slow and fast dynamic scatterer components to the flow index derived by continuous wave diffuse correlation spectroscopy

A light-based blood-flow sensor gets confused by 'boring' tissue — here's how much it skews results.

Doctors increasingly use a laser-light technique called diffuse correlation spectroscopy to non-invasively track blood flow in tissue, by shining light in and watching how the pattern flickers as red blood cells move. But tissue isn't just moving blood cells — it also contains completely still material and slowly-shifting material, both of which distort the flicker pattern and can throw off the blood-flow reading. This study built physical models ('phantoms' — like gel blocks with a tube of flowing liquid standing in for blood vessels) to measure exactly how much these non-blood components skew the results. Understanding and correcting for this matters because it affects how accurately doctors can trust these devices for monitoring things like brain or muscle blood flow in real patients.

Technical view

Using continuous-wave diffuse correlation spectroscopy (cw-DCS), the authors quantify how static and slow-dynamic scatterer populations bias the derived blood flow index (BFI), which is conventionally computed assuming decorrelation is dominated solely by fast-moving red blood cells. Agar-based phantoms embedded with a flow tube were used to isolate and measure the fractional contribution of each scatterer class (static, slow-dynamic, fast-dynamic) to the measured autocorrelation decay rate. Results show static/slow components meaningfully alter derived BFI values, implying that current single-component fitting models used in clinical cw-DCS devices may need multi-component correction terms for accurate absolute (not just relative) flow quantification — directly relevant to anyone calibrating or validating DCS hardware.

bioRxiv · bioinformaticsRunnable

Benchmark Averages Hide the Failures That Matter: Quantizing ESM-2 for Protein Variant-Effect Prediction

Shrinking a protein-AI model to save memory can secretly break it on the one case that matters.

Big AI models that predict how mutations affect proteins are often 'compressed' (quantized) to run faster and cheaper, using lower-precision numbers instead of full precision, and this is usually judged safe by checking if average accuracy across many tests barely moves. This study tested six different compression settings on a popular protein-AI model (ESM-2) across 201 real mutation-effect benchmarks covering 2.4 million variants. The catch: even when the average accuracy looked essentially unchanged, one specific compression method silently wrecked performance on an individual test, cutting its accuracy score by more than half. This means relying on average benchmark scores to decide if a compressed AI model is 'safe' to deploy can hide serious, deployment-breaking failures on important individual cases.

Technical view

The authors benchmark six numerical precision configurations of ESM-2 (650M–15B parameters) on bulk embedding extraction and deep mutational scanning (DMS) variant-effect scoring, evaluated across the full ProteinGym substitution benchmark (201 assays, 2.41M variants) with paired bootstrap clustering by protein. While no configuration shifts mean Spearman correlation by more than 0.007 at any scale, INT8 dynamic quantization — statistically indistinguishable from fp32 on the mean at 3B parameters (p=0.34) — collapses one specific assay's correlation from ρ=0.591 to 0.223. This demonstrates that mean-based benchmark evaluation is an unreliable criterion for quantization deployment decisions in protein language models, and practitioners should instead audit worst-case, per-assay degradation before shipping quantized variant-effect predictors.

bioRxiv · biophysicsBuildable

De novo design of autocatalytically forming intra- and intermolecular isopeptide bonds to construct rigid covalent protein assemblies

Scientists made proteins glue themselves together permanently, no enzymes needed.

Some bacteria build super-tough surface fibers by having two amino acids in a protein spontaneously fuse into a permanent chemical bond, no external glue or enzyme required. Here researchers designed brand-new proteins from scratch that do this same self-stitching trick, both within a single protein chain and between two separate ones. They even split the design into two halves that only snap together and lock when mixed, and tuned it so temperature controls exactly when the bond forms. The payoff is being able to build large, rigid, precisely shaped molecular structures — like a 215,000-atomic-mass-unit ring — that are essentially welded together and won't fall apart, useful for building durable nanoscale materials or vaccine scaffolds.

Technical view

The authors computationally designed de novo proteins that autocatalytically form intramolecular and intermolecular isopeptide bonds (side-chain amide linkages, as seen in Gram-positive pilin proteins), producing 50+ validated designs confirmed by mass spec and 5 crystal structures. They engineered split constructs that crosslink upon combination, orthogonal to the widely used SpyTag/SpyCatcher system, with bond formation kinetics tunable by temperature. These modules were used to rigidly crosslink domains into large (up to 215 kDa) symmetric, covalently locked ring assemblies — a generalizable toolkit for programmable, irreversible protein nanoarchitecture beyond existing peptide-tag crosslinking systems.

bioRxiv · cancer biologyConceptual

Tyrosine phosphorylation and dimerization cooperatively activate NAMPT to enable NAD+ synthesis in cancer

Cancer cells hijack a metabolic enzyme by tagging and pairing it up for extra fuel.

Cells need a molecule called NAD+ to run their metabolism, and an enzyme called NAMPT is the key factory that makes it. This study found that several cancer-driving proteins (kinases that are often mutated or overactive in tumors) directly chemically tag NAMPT at one specific spot, and this tag works together with NAMPT pairing up with a copy of itself to switch the enzyme into overdrive. When researchers blocked either the tagging or the pairing, cancer cells made less NAD+, grew slower, and formed fewer colonies. This matters because it reveals a druggable weak point — cutting off this tag-and-pair switch could starve cancer cells of the fuel-making machinery they depend on.

Technical view

NAMPT, the rate-limiting NAD+ salvage-pathway enzyme, is shown to be a direct phosphorylation substrate of oncogenic tyrosine kinases (ALK, INSR, IGF1R, PDGFRA), with phosphoproteomics identifying Y188 as the dominant site, notably downstream of the NPM1::ALK fusion. Y188 phosphorylation and NAMPT dimerization act cooperatively to boost catalytic activity and NMN/NAD biosynthesis; a Y188F mutant or dimerization-disrupting mutation each independently impair enzymatic activity, proliferation, and clonogenic growth. This defines a kinase-NAMPT signaling axis as a candidate therapeutic target for NAD+-dependent tumors, with Y188 phosphorylation or dimer-interface disruption as potential intervention points.

bioRxiv · cancer biologyBuildable

Deep learning-based identification and quantification of rare circulating hybrid cells in orthotopic pancreatic cancer models

AI microscope hunts for one weird hybrid cancer cell among millions of normal blood cells.

When cancer spreads, some rare cells appear to be hybrids — part tumor, part immune cell — floating in the bloodstream, and finding these needle-in-a-haystack cells under a microscope is extremely hard because they're so sparse and every sample looks different. The researchers built a two-step system: first they use each animal's own unstained blood sample as a personalized baseline to filter out background noise, then they train a neural network (a pattern-recognizing AI) on labeled cell images to flag genuine candidates. They tested this in mice with pancreatic tumors, imaging blood cells tagged with fluorescent markers for tumor and immune cell identity. This kind of automated detection could make it far more practical to track cancer spread in blood samples instead of painful tissue biopsies.

Technical view

The authors present a two-stage pipeline for detecting rare ECAD+/CD45+ circulating hybrid cells (CHCs) in mouse PBMC preparations via multichannel fluorescence microscopy: first, animal-specific unstained control images establish per-mouse background distributions for candidate enrichment, then a CNN trained on blinded multi-annotator consensus labels (DAPI, ECAD, CD45 crops) classifies candidates. Validated across 28 mice (tumor-bearing and tumor-naive), this specimen-normalized enrichment-plus-classification approach addresses the core challenge of rare-cell detection amid inter-sample background variability, and the framework could generalize to other rare circulating cell types with adaptation of channel inputs and training labels.

bioRxiv · cancer biologyRunnable

PDE3A-SLFN12 Molecular Glues Target Multiple KIT D816V Cell Types in Preclinical Models of Mast Cell Malignancies

A newly found drug tricks two proteins into teaming up to kill mutant mast cell cancer.

Certain mast cell cancers are driven almost entirely by one specific mutation in a gene called KIT, and this study looked for a drug that kills only cells carrying that mutation while sparing normal cells. Using cells grown from stem cells of actual patients, the team screened a large library of drugs and found one, called LDC3416, that selectively kills the mutant cells. Digging into how it works, they discovered it acts as a 'molecular glue' — essentially super-gluing two proteins (PDE3A and SLFN12) together into a cell-killing complex, and the mutant cancer cells happen to make extra amounts of both proteins, making them uniquely vulnerable to this glue effect. This offers a promising, more targeted treatment strategy for a disease that currently has few options.

Technical view

Using KIT D816V patient-iPSC-derived hematopoietic and mast cell models, the authors screened FDA-approved and experimental compounds and identified LDC3416 as selectively cytotoxic to KIT D816V-mutant cells across multiple lineages. Mechanistic profiling revealed LDC3416 acts through the PDE3A-SLFN12 molecular glue pathway, and that KIT D816V signaling upregulates both PDE3A and SLFN12 expression, conferring selective sensitivity to glue-induced complex formation and cell death. This establishes molecular-glue induction of the PDE3A-SLFN12 axis as a mutation-selective therapeutic strategy for clonal mast cell malignancies, with LDC3416 as a lead compound for further preclinical development.

bioRxiv · systems biologyConceptual

Potential benefit of loss-of-function on bacterial fitness

Losing genes can actually help bacteria — if the gene was expensive to keep making.

Every gene a bacterium carries costs it energy and resources to constantly manufacture into protein, so scientists wondered whether ditching some genes could actually free up resources and boost growth. Using E. coli as the test subject, they sorted genes into categories — essential, important, mostly-neutral, or even fitness-boosting when deleted — and combined this with data on how much of the cell's protein-making budget each gene consumes. They found that expensive-to-produce genes are more likely to matter for fitness, but surprisingly, deleting genes rarely helps growth simply by freeing up that production budget. This nuances the popular idea that bacteria are constantly trimming costly, useless genes to save energy — the real picture of why gene loss helps is more complicated.

Technical view

The study integrates fitness measurements with proteomic abundance data across the E. coli genome to test whether gene loss improves fitness primarily by reducing proteome allocation cost. Genes are classified into essential, important, mean-effect, and fitness-enhancing categories; mean-effect genes constitute 31-75% of the total proteome mass fraction depending on growth condition (highest in LB medium), with core mean-effect genes enriched for transmembrane transport functions. Comparison against genome-scale metabolic models with proteome constraints (ME-models) shows that while high proteomic-cost genes are more likely to influence fitness, fitness-enhancing deletions rarely act via simple resource-reallocation savings — implying deletion benefits arise from more complex regulatory or metabolic effects than proteome-burden relief alone.

bioRxiv · neuroscienceConceptual

Diurnal time and sleep pressure modulate cerebrospinal fluid low-frequency oscillations in synchrony with wake-promoting nuclei

Staying awake longer makes your brain fluid pulse harder to flush out waste.

While you sleep, slow rhythmic pulses in your brain's blood vessels help pump cerebrospinal fluid (the liquid cushioning your brain) through brain tissue, flushing out waste — a process called brain clearance. This study kept healthy people awake for 34 hours straight while continuously scanning their brains and measuring alertness, to see what drives these fluid pulses beyond just sleep itself. They found the pulses got stronger the longer people stayed awake and also naturally peaked in the early morning regardless of sleep loss, and this tracked closely with alertness levels and activity in brain circuits that keep us awake — but not with a stress hormone or with typical sleep-related brain wave patterns. This suggests the brain's built-in waking and body-clock systems, not just sleep itself, help regulate how well it cleans itself, which matters for understanding conditions linked to poor sleep and waste buildup, like some neurodegenerative diseases.

Technical view

In a 34-hour sleep deprivation protocol with dense longitudinal fMRI/EEG and NIRS sampling in healthy adults, the authors quantified cerebrospinal fluid low-frequency oscillations (LFOs, the vasomotion-driven signal thought to propel CSF-based brain clearance) and dissociated circadian versus homeostatic sleep-pressure contributions. CSF LFO amplitude increased monotonically with time awake and showed independent diurnal modulation peaking in early morning; LFOs correlated strongly with vigilance measures and reticular activating system (wake-promoting brainstem/thalamic) activity, but not with plasma norepinephrine or EEG slow-wave activity. The findings implicate arousal-circuit and circadian signaling, rather than adrenergic tone or classic sleep EEG markers, as regulators of vasomotor-driven CSF dynamics — relevant to models linking sleep loss, arousal state, and glymphatic-type clearance dysfunction.

bioRxiv · molecular biologyConceptual

Healing of chromosomal breaks is impeded in cells expressing progerin

A protein tied to premature-aging disease jams the cell's DNA repair crew.

Hutchinson-Gilford Progeria Syndrome is a devastating rare disease that makes children age dramatically fast and die young, caused by a mutation that produces a toxic, truncated protein called progerin (a warped version of a normal structural protein in the cell's nucleus). Cells with progerin build up broken DNA strands and seem worse at fixing them, so this study investigates exactly how progerin interferes with the two main repair systems cells use to mend broken DNA — one precise (using a matching template) and one quicker but sloppier. Using cells engineered to make progerin, the researchers show that break repair is impaired, and start dissecting which repair pathway suffers and why. Understanding this could explain how progerin drives accelerated aging and possibly shed light on ordinary aging too, since everyone makes small amounts of progerin naturally.

Technical view

HGPS results from a LMNA point mutation activating a cryptic splice site, producing farnesylated truncated lamin A ('progerin'), which is associated with elevated genomic double-strand break (DSB) burden and altered DSB repair balance between homologous recombination (HR, template-directed and accurate) and end-joining (EJ, error-prone). The authors use progerin-expressing cell models to directly assess DSB repair kinetics and pathway usage, finding that break healing is impeded in progerin-expressing cells. This work aims to mechanistically link progerin expression to specific repair pathway defects, with implications for both HGPS pathology and physiological aging, since low-level progerin production occurs in normal cells as well.

bioRxiv · molecular biologyBuildable

Multiplexed construction of defined DNA products from unpartitioned oligonucleotide pools

A new method builds dozens of custom DNA sequences at once from one cheap mixed batch.

Making custom DNA in the lab is expensive when done piece by piece, but far cheaper when you order many short DNA snippets together as a mixed 'pool' — the catch is sorting that jumbled pool back into the exact separate DNA products you actually want. This paper introduces a method called Oligo Pool Sidewinder assembly that solves this sorting problem: it uses a computer program to design each short snippet with built-in matching rules, so that in a single test tube reaction, hundreds of pieces spontaneously find their correct partners and stitch together correctly into dozens of distinct final DNA constructs. This makes large-scale, low-cost DNA construction from cheap pooled snippets practical, which matters for anyone building genes, gene circuits, or other custom DNA at scale — from researchers to biotech companies.

Technical view

The paper introduces Oligo Pool Sidewinder assembly, a method combining a computational bespoke-oligo design workflow with a defined set of construction rules to enable one-pot, parallel, highly multiplexed assembly of hundreds of DNA fragments from unpartitioned synthetic oligo pools into dozens of distinct, correctly-formed double-stranded DNA constructs simultaneously. This directly addresses the yield/fidelity tradeoff inherent to pooled oligo synthesis (cheap but unindexed) versus individually synthesized oligos (isolatable but costly), offering a scalable route to reduce cost, labor, and turnaround for de novo multi-construct DNA production, with the computational design component likely adaptable to other multiplexed assembly schemes.

bioRxiv · cell biologyBuildable

Single-Cell Mapping of tRNA Expression Dynamics Across Human Hematopoiesis

A new lab trick reveals how blood stem cells retune their protein-building machinery as they mature.

Every cell uses molecules called tRNAs to translate genetic instructions into proteins, like adapters that match each three-letter DNA code to the right amino-acid building block. Scientists didn't know if these adapters change as a blood stem cell decides to become, say, a red blood cell or an immune cell. Here researchers built a new tool that reads both a cell's tRNAs and its regular gene activity at the same time, in thousands of individual cells from human bone marrow. They found specific tRNA patterns line up with 'stemness' versus different mature blood cell fates, revealing a whole extra layer of cellular identity nobody had mapped before. This matters because it suggests cells fine-tune protein production, not just gene activity, to become who they are.

Technical view

The authors developed sc-STM-seq, a scalable single-cell method that co-profiles tRNA abundance/splicing and mRNA transcriptomes in the same cell, addressing the technical difficulty of capturing short, heavily modified tRNAs alongside standard scRNA-seq. Applied to human bone marrow, they construct differentiation trajectories and correlate specific tRNA isodecoder expression and splicing events with stemness and lineage commitment. The result is a hematopoietic tRNA atlas positioning tRNA regulation as an independent axis of cellular heterogeneity beyond transcript-level control. Practitioners in translational control or codon-usage optimization could apply sc-STM-seq to other differentiation systems or mine this dataset for lineage-specific tRNA biomarkers.

bioRxiv · cell biologyBuildable

Trustworthy in silico labeling via semantic visual interpretability of image-to-image translation

A new tool cracks open the black box of AI that paints fake microscope images of cell organelles.

In silico labeling is a trick where an AI looks at a plain, unstained microscope image and predicts where specific cell structures like mitochondria would glow if you'd stained them, saving time and chemicals. The problem is nobody could tell if the AI was truly 'seeing' real biology or just guessing from unrelated visual patterns, which makes it risky to trust. The researchers built Mask Interpreter, a method that shows exactly which visual features the AI relies on for each organelle, like a fingerprint of its reasoning. It turns out the models genuinely key off real biological structures, and the tool can also catch experimental glitches ('batch effects') and flag likely mistakes even without a ground-truth image to compare against. This matters because it makes a powerful but opaque prediction technology safe enough to actually trust in biology labs.

Technical view

Mask Interpreter is a semantic visual interpretability framework for image-to-image translation models used in label-free organelle prediction, generating organelle-specific 'explanation signatures' rather than generic saliency maps. Benchmarked against standard explainable-AI (xAI) methods, it better discriminates authentic biological signal from spurious correlations the network may exploit, and additionally functions as a diagnostic for batch effects and localized prediction error even absent paired fluorescence ground truth. This positions it as both a validation layer for deploying in silico labeling pipelines and a QC tool for microscopy datasets. Groups building or auditing cross-modality translation models could adopt the explanation-signature approach as a drop-in interpretability check.

bioRxiv · cell biologyConceptual

Molecular orchestration of global actin remodelling by INF2

Cells have a fast-acting 'reset button' for their internal skeleton, and scientists just found its safety switch.

Cells contain a scaffold of protein filaments called actin that reshapes itself constantly to change a cell's shape or heal a wound. When calcium floods into a cell, a protein called INF2 triggers a rapid, sweeping rebuild of this scaffold, but too much INF2 activity is linked to kidney and nerve diseases, so it needs strict control. Using high-resolution microscopy that tracks single molecules plus biochemistry and structural analysis, the researchers found two built-in brakes: one where INF2 folds up on itself, and another where its tip grabs the side of actin filaments to stop them growing too long. Breaking this second brake keeps INF2 switched on too long and disrupts the cell's outer membrane. This matters because it explains, at a molecular level, how cells prevent a normal repair process from spiraling into disease.

Technical view

The study dissects the regulatory circuit governing INF2, the formin responsible for the Calcium-mediated Actin Reset (CaAR) reaction, combining live-cell imaging, single-molecule tracking, biochemistry, and structural analysis. Beyond the known intramolecular autoinhibition, they identify a second mechanism in which the INF2 N-terminus binds laterally to actin filament sides, capping elongation and promoting reversion to the autoinhibited state, a negative feedback loop. Disrupting this side-binding interaction prolongs INF2 activity and perturbs plasma membrane organization, linking the mechanism directly to the kidney and neuronal pathologies associated with INF2 hyperactivity. This gives structural and cell biologists a specific interface (N-terminus-filament side contact) as a candidate target for modulating formin activity therapeutically.

bioRxiv · biochemistryBuildable

Activity-based chemical proteomics uncovers unexpected covalent targetsof E64d and reveals a role for cysteine cathepsins in PLD3 proteostasis.

A drug tool meant to study one enzyme turned out to be hitting an Alzheimer's-linked protein too.

PLD3 is a protein tied to Alzheimer's disease and immune signaling, but it only works once other enzymes chop it into its active form, and nobody knew which enzymes did the chopping. Researchers used a chemical called E64d, known to block a family of protein-cutting enzymes called cysteine cathepsins, to investigate. To make sure E64d was doing what everyone assumed, they built a modified 'tagged' version of it and used a technique that maps every protein a drug actually sticks to in a cell. Surprisingly, E64d also grabbed onto several other proteins beyond cathepsins, including ones linked to detox pathways and gene expression, while also pointing to which cathepsins process PLD3. This matters because it both clarifies a mechanism relevant to Alzheimer's and warns researchers to double-check what their 'specific' inhibitors are really doing.

Technical view

The authors use activity-based protein profiling (ABPP) with a propargyl-alkyne analogue of the cysteine cathepsin inhibitor E64d to map its full covalent target landscape in cells, going beyond its canonical cathepsin targets. The chemoproteomic survey uncovers unexpected covalent engagement with bleomycin hydrolase (BLMH), KEAP1, SPT5 (SUPT5H), and an asparagine synthetase, indicating substantial off-target reactivity for a widely used cathepsin probe. In parallel, they use this pharmacological toolkit to implicate specific cysteine cathepsins in the proteolytic maturation of PLD3, a 5'-3' exonuclease relevant to Alzheimer's disease and innate immune signaling. This is directly useful to chemical biologists as a target-engagement dataset for reinterpreting past E64d-based studies and for cathepsin-selective probe design.

bioRxiv · biochemistryConceptual

Structural basis of K<+>/H<+> antiport in YcgO and its inhibition by unphosphorylated PtsN

Cryo-electron microscopy caught a bacterial potassium pump in the act of being switched off.

Cells need to carefully balance potassium and hydrogen ions to survive, a job done by proteins called cation-proton antiporters that swap one ion for the other across the cell membrane. Researchers studied one such pump in E. coli, called YcgO, which trades potassium for protons, and a regulatory protein called PtsN that can shut it down. Using cryo-electron microscopy, a technique that flash-freezes molecules and images them with electrons to reveal their 3D shape, they captured detailed snapshots of YcgO both alone and locked by PtsN. The structures show YcgO is a two-part (dimeric) machine with extra internal 'sensor' domains that PtsN grabs to jam the pump's moving parts. This matters because it reveals a previously unclear architecture and control mechanism for a whole class of potassium transporters found across many organisms.

Technical view

Using cryo-EM at 3.4 Å and 3.2 Å resolution, the authors solve structures of the E. coli CPA1 antiporter YcgO in its K+-bound occluded state and in complex with the unphosphorylated regulatory protein PtsN. YcgO is a homodimer with appended cytosolic RCK and CorC domains that couple ion-binding events to conformational changes in the transport helices; PtsN docks onto the CorC domains to allosterically inhibit helix movement and thus efflux. This provides a structural rationale for phosphorylation-state-dependent regulation of a K+-specific CPA1 transporter via a PTS-system component, a regulatory logic distinct from previously characterized Na+/H+ antiporters. Structural biologists working on ion transport or bacterial signaling could use these coordinates to model analogous RCK/CorC-domain regulatory interfaces in other CPA-family transporters.

Q

Quanta — Explained

1 new
Quanta MagazineConceptual★ flagship

Theory of Fluids Enters the 21st Century

Physicists rewrite the 200-year-old rulebook for how liquids behave, from the atoms up.

For roughly two centuries, physicists described fluids—how liquids flow, mix, and settle—using essentially one classic theory that treats the liquid as a smooth continuous stuff. The trouble is that liquids are actually made of jostling molecules, and the old picture glosses over how that molecular graininess shapes everyday behavior. Using a modern theoretical insight, researchers rebuilt the description starting from the individual particles and their interactions, deriving fluid behavior 'from the bottom up' rather than assuming it. The payoff is a more fundamental and accurate account of when and why the old theory works or breaks down, which could sharpen how we model everything from biological cells to industrial processes.

Technical view

This is a Quanta Magazine feature (not a primary paper), so specifics are limited, but it reports a reformulation of liquid-state theory grounded in a modern insight that rederives fluid behavior from microscopic particle interactions rather than the long-standing continuum framework. The claim is a bottom-up redefinition that clarifies the domain of validity of classical fluid theory. A technical reader should treat this as a pointer: follow the cited researchers and underlying publications to find the actual formalism, assumptions, and testable predictions before building on it.

HN

What's Trending

60 new
Hacker News · 1061 ptsConceptual★ flagship

AI;DR (AI; Didn't Read)

A riff on 'TL;DR'—when AI does the reading, and skimming, for you.

The title plays on 'TL;DR' (too long; didn't read), swapping in AI to suggest a world where an AI reads and digests content on your behalf. The underlying theme is how AI summarizers increasingly stand between people and the original material, so many of us consume a machine's condensed version instead of the source. That raises real questions about what gets lost, distorted, or quietly editorialized when a model decides what matters. Without an abstract the specifics are unclear, but the piece appears to poke at this shift in how we read and trust information in an AI-mediated world.

Technical view

No abstract or body is provided, so the substance can't be pinned down beyond the wordplay ('AI; Didn't Read' as an AI-driven successor to TL;DR). Thematically it points at automated summarization and the epistemic effects of consuming model-generated digests rather than primary sources. A technical reader interested in this space would look toward summarization faithfulness, hallucination in condensation, and information-loss or bias metrics—but nothing here specifies a method, dataset, or result to replicate.

Hacker News · 925 ptsConceptual★ flagship

The Amazon tax

The hidden cost of routing so much of the economy through one giant platform.

The phrase 'the Amazon tax' typically refers to the effective cost—paid by sellers, and often passed on to shoppers—of doing business through a dominant marketplace that takes fees, ads, and logistics cuts on every sale. The idea is that when one platform becomes the default way to reach customers, participating stops being optional and its rising charges act like a tax on commerce. No abstract is provided here, so the exact argument isn't specified, but the framing usually explores how platform power translates into pricing, seller margins, and competition. It matters because these dynamics shape what products exist, what they cost, and who captures the value in online retail.

Technical view

With no abstract supplied, the specific thesis can't be verified; the title is a common shorthand for the aggregate take-rate a dominant marketplace extracts (referral fees, fulfillment, advertising) that behaves economically like a tax on sellers. Analytically this connects to platform economics, market power, and vertical relationships between a marketplace operator and third-party sellers. A rigorous treatment would quantify effective take rates over time, pass-through to consumer prices, and margin compression for sellers—but none of those figures or methods are given here to build on.

Hacker News · 811 ptsConceptual

Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

An op-ed argues Claude secretly tweaks your prose to leave a hidden AI fingerprint, and that's cheating writers.

Some AI systems reportedly embed subtle, detectable patterns into the text they generate, like unusual word choices or phrasing quirks, so the output can later be identified as AI-written, a technique often called watermarking. This piece argues that when Claude does something like this, it's quietly altering the writing itself rather than just adding an invisible tag, which the author calls a 'perversion' of what writing is supposed to be. The underlying worry is that a tool meant to help you write is instead prioritizing its own traceability over giving you the best possible sentence. Since only the title is available here, treat this as a critical opinion rather than a confirmed technical report.

Technical view

The piece appears to be commentary criticizing an alleged text-watermarking behavior in Claude's outputs, where token-level or stylistic biasing meant to make AI-generated text detectable is claimed to measurably distort the writing itself. Without further detail than the title, the specifics (mechanism, detection method, or whether this is an official Anthropic feature versus inferred behavior) aren't established here. Readers interested in AI-text watermarking mechanics generally should look at published statistical watermarking schemes (e.g., green-list token biasing) to understand the detectability-versus-quality tradeoff this critique is gesturing at.

Hacker News · 784 ptsRunnable

Qwen 3.8 27B is excellent, but it defaults to overthinking things

This open AI model is smart but can't stop double- and triple-checking simple answers.

Qwen 3.8 27B is a large language model, a type of AI trained to understand and generate text, that reviewers are praising as genuinely strong. Its one big flaw is 'overthinking': like other reasoning-style models, it often generates long chains of internal step-by-step deliberation even for questions that don't need it, making responses slower and more bloated than necessary. This matters to anyone choosing an AI model because raw capability isn't the whole story; how efficiently a model uses its 'thinking' time affects cost, speed, and usability in real products.

Technical view

Qwen 3.8 27B is reported as a capable model in the Qwen family but exhibits excessive test-time reasoning ('overthinking'), generating unnecessarily long chain-of-thought traces even on low-difficulty prompts, a known failure mode in reasoning-tuned LLMs that increases latency and token cost without proportional accuracy gains. Practitioners evaluating it should benchmark reasoning-token length versus task difficulty and consider prompting strategies or reasoning-budget controls to mitigate the overhead. As an open-weight model, it can be run and profiled locally to characterize this behavior directly.

Hacker News · 707 ptsConceptual

Universal health coverage could save $1T and 114k lives a year: study

Giving everyone basic health coverage could save a trillion dollars and over 100,000 lives yearly.

Universal health coverage means everyone in a country can get basic medical care without being pushed into poverty by the cost. A new study estimates that expanding this kind of coverage globally wouldn't just save lives, around 114,000 per year, it would also save roughly $1 trillion annually, likely by preventing costlier emergency care and lost productivity from untreated illness. This reframes health coverage as an investment that pays for itself rather than purely a cost, which matters for governments deciding how to spend limited budgets. Since only the headline figures are given here, the underlying methodology and country scope aren't detailed.

Technical view

The study models the economic and mortality impact of scaling universal health coverage (UHC), projecting roughly $1 trillion in annual savings alongside ~114,000 averted deaths per year, though the available summary doesn't specify the modeling framework, cost categories, or geographic scope. Such estimates typically combine averted direct treatment costs, productivity gains from a healthier workforce, and value-of-statistical-life calculations for the mortality reduction. Health economists or policy researchers would need the full report to assess assumptions (discount rates, coverage baseline, cost sources) before using these figures in comparative policy analysis.

Hacker News · 702 ptsConceptual

A Preview of DuckDB v2.0

The database that fits in your pocket is about to get a major upgrade.

DuckDB is a lightweight database that lives inside your own program instead of running as a separate server — think of it as a supercharged spreadsheet engine you can embed anywhere, from a laptop script to a data pipeline. It's built for quickly crunching large datasets (millions of rows) with SQL, the standard language for asking questions of data. Version 2.0 is a major milestone release, the kind that usually brings faster performance, new capabilities, and possibly some changes to how existing features work. It matters because DuckDB has become a go-to tool for data scientists and engineers who want fast analysis without the hassle of setting up a full database server.

Technical view

DuckDB is an embedded OLAP (online analytical processing) engine using a columnar, vectorized execution model, positioned as 'SQLite for analytics.' A v2.0 release signals a major-version milestone, typically bundling breaking API/storage changes, performance work on the vectorized execution engine, and expanded SQL/extension support. Practitioners tracking DuckDB for ETL, embedded analytics, or as a Pandas/Polars alternative should watch for storage format compatibility and extension ecosystem changes before upgrading production pipelines.

Hacker News · 683 ptsBuildable

How Bluesky draws its logo on screenshots

Bluesky quietly stamps its logo onto screenshots the moment you try to capture one.

Bluesky is a social network similar to Twitter/X, and this piece explains a clever trick their app uses: when you take a screenshot of a post, the app detects that action and overlays its logo onto the image before it gets saved. The real-world problem this solves is branding and attribution — when screenshots of posts get shared around the internet (which happens constantly with social media), the origin of that content can get lost. Their approach involves hooking into the device's screenshot-detection APIs and dynamically rendering a watermark in that split second. It matters as a neat example of using platform-level APIs creatively to solve a branding problem without changing the actual user experience of browsing.

Technical view

The piece reverse-engineers or documents how Bluesky's client listens for OS-level screenshot events (e.g., UIApplication.userDidTakeScreenshotNotification on iOS, or equivalent Android broadcast receivers) and injects a rendered logo overlay into the captured view hierarchy at the moment of capture. This is a practical technique for any app wanting to brand user-generated screenshots without persistent on-screen watermarks. Developers could replicate this with native screenshot-detection hooks paired with a just-in-time view redraw or compositing layer.

Hacker News · 628 ptsConceptual

Ask HN: Alternatives to GitHub

GitHub keeps going down, so developers are asking: is it time to jump ship?

This is a discussion thread on Hacker News, a tech forum, where someone points out that GitHub — the massive code-hosting platform millions of developers rely on daily — has been having repeated outages lately. The question posed is simple but weighty: should developers and teams start moving their projects to alternative platforms instead of waiting out GitHub's reliability issues? The 'approach' here isn't a study but a crowd-sourced conversation, where people will likely list options like GitLab, Codeberg, or self-hosted tools, and weigh trade-offs like ease of migration, feature parity, and community size. It matters because so much of the software world's infrastructure — code, collaboration, CI/CD pipelines — depends on GitHub staying up, so recurring outages raise real risk-management questions for teams and companies.

Technical view

This is a community discussion thread (Ask HN format) rather than a technical report, prompted by observed reliability degradation in GitHub's uptime over recent months. Expect crowd-sourced comparisons of alternatives such as GitLab (self-hosted or SaaS), Codeberg, sourcehut, Gitea, and Forgejo, likely weighing migration costs (issues, PRs, CI config, webhooks) against redundancy benefits. Anyone evaluating platform risk could use this thread as an informal survey of real-world alternative-hosting experiences and migration pain points.

Hacker News · 617 ptsRunnable

GPT-5.6 Sol Pricing Cut by 50% on OpenRouter

A cutting-edge AI model just got half-price overnight on a model marketplace.

OpenRouter is a service that lets developers access many different AI models through one unified interface, kind of like a price-comparison site for AI. GPT-5.6 Sol is one such AI language model, and its price for developers to use just dropped by 50% on that marketplace. This kind of price cut usually happens because of increased competition among AI providers, improved efficiency in running the model, or a push to get more developers building on top of it. It matters to anyone building AI-powered apps because the cost of calling these models directly affects what products are financially viable to build.

Technical view

GPT-5.6 Sol's per-token pricing was cut by 50% on OpenRouter, an API aggregator/router that lets developers switch between LLM providers via a unified endpoint. Such cuts typically reflect either provider-side price competition, improved inference efficiency (e.g., quantization, better batching), or promotional positioning against rival models. Developers building on OpenRouter can take advantage immediately by routing existing calls to the model without code changes beyond the model identifier, making it a low-friction cost optimization to test in production.

Hacker News · 565 ptsConceptual

Google has acquired the data of failed US airline Spirit

When an airline goes bankrupt, its passenger data can become someone else's asset — this time, Google's.

Spirit Airlines, a budget US airline, went out of business, and as part of that failure, its data assets — likely including customer information like flight histories, loyalty program details, or contact info — were acquired by Google. This happens because when a company fails, its assets, including data, get sold off during bankruptcy proceedings, similar to how furniture or equipment would be auctioned. The 'approach' here is essentially a corporate acquisition or licensing deal rather than any new technology. It matters because it raises questions about what happens to your personal data when a company you trusted goes bankrupt — it can end up in the hands of a completely different company, like a tech giant, without much say from the original customers.

Technical view

Following Spirit Airlines' business failure, Google reportedly acquired its data assets — likely passenger records, booking history, or loyalty data — through bankruptcy asset liquidation proceedings. This raises data governance and privacy questions: bankrupt companies' data assets are typically sold under court-supervised sale processes, and buyer obligations around the original privacy policies vary by jurisdiction and sale terms. Anyone tracking data privacy law or corporate M&A should watch for details on what specific data was included and what consent/opt-out mechanisms, if any, applied to affected Spirit customers.

Hacker News · 520 ptsBuildable

Olo (Color)

A new tool called Olo wants to change how you think about and pick colors.

Olo appears to be a project or tool centered on color — likely a color picker, color space, or design tool for developers and designers. Without more detail, the core idea seems to be giving people a better or more intuitive way to work with color, which is a surprisingly tricky problem: colors look different on different screens, and picking harmonious color combinations requires either expertise or good tooling. The approach is presumably some kind of interactive or algorithmic system for selecting, converting, or reasoning about color values. It would matter to designers, developers, and anyone building visual products who wants more control or clarity when working with color.

Technical view

Olo (Color) appears to be a tool or library focused on color manipulation, selection, or a color space/model, though the abstract doesn't specify the exact mechanism. Given the 'Show HN'-style title format, it's likely a small utility or library that developers could integrate directly — for example, a perceptually uniform color space converter, a palette generator, or a color picker UI component. Practitioners interested in color science or design tooling should check the source for specifics on the underlying color model (e.g., OKLCH, LAB) before adopting it.

Hacker News · 507 ptsRunnable

Linux 7.3 improves performance when running out of vRAM

Linux just got smarter about not choking when your graphics card runs out of memory.

Your computer's graphics card (GPU) has its own dedicated memory called vRAM, used to store the visuals and data needed for rendering images, games, or AI workloads. When a task needs more vRAM than the GPU actually has, the system has to improvise — often by spilling data over to slower regular system memory, which can cause slowdowns or stutters. This new version of the Linux operating system kernel includes improvements to handle that 'running out of vRAM' situation more gracefully and with better performance. It matters to gamers, AI researchers, and anyone doing graphics-heavy work on Linux, since running out of vRAM is a common bottleneck, especially with today's memory-hungry AI models.

Technical view

Linux 7.3 includes kernel-level improvements to GPU memory management for scenarios where vRAM is exhausted and the system must fall back to system RAM (via mechanisms like TTM/GEM memory eviction or unified memory paths in drivers such as amdgpu or the DRM subsystem). Better performance under vRAM pressure typically comes from smarter eviction policies, reduced overhead in the fallback path, or improved memory oversubscription handling. This is directly relevant to workloads that regularly exceed GPU memory capacity, such as large AI model inference/training or high-resolution rendering, where practitioners can expect fewer stalls or OOM-related crashes.

Hacker News · 479 ptsConceptual

Memory prices climb 500% in 12 months

The RAM chips inside your devices are now six times pricier than they were a year ago.

Memory chips — the RAM that lets your computer and phone quickly store and access data while running programs — have shot up in price by 500% over the past year, an enormous jump for essential computer hardware. This kind of price surge usually happens when demand suddenly outpaces supply, and the massive boom in AI has been driving huge demand for memory chips used in data centers and AI training hardware, squeezing supply for everyone else. It matters because this price spike ripples through the entire tech industry, making laptops, phones, and other gadgets more expensive to build, and it's a very visible sign of how much AI infrastructure is reshaping the broader hardware economy.

Technical view

DRAM (and likely NAND) spot and contract prices have risen roughly 500% year-over-year, driven primarily by AI datacenter demand for high-bandwidth memory (HBM) and general DRAM consuming fab capacity that would otherwise serve consumer and enterprise markets. This supply-demand imbalance is squeezing memory-dependent industries — device OEMs, server manufacturers, and consumer electronics — and is likely to affect BOM (bill of materials) costs across the hardware stack through at least the next few product cycles. Engineers and product teams planning hardware releases should factor in elevated memory costs and potential allocation constraints when sourcing components.

Hacker News · 477 ptsConceptual

Quake Shareware, a CD-ROM just a little too full

How id Software squeezed the entire Quake demo onto a nearly-full 90s CD-ROM.

Back in the 1990s, games were often distributed as 'shareware' — a free trial version handed out on CD-ROM to hook players before they bought the full game. This piece digs into how the Quake shareware disc was assembled, and how close the developers cut it to the CD's actual storage limit. It walks through the constraints of the era: CDs held a fixed, tiny amount of data by today's standards, and every texture, sound, and level had to be trimmed or compressed to fit. It's a fun window into the physical, practical side of old game development, when running out of disc space could mean cutting content.

Technical view

The piece is a technical retrospective/forensic look at the original Quake shareware CD-ROM image, examining how id Software packed assets (levels, textures, sound) to fit within the ~650-700MB capacity of a standard CD, likely including filesystem layout, compression choices, and near-capacity byte accounting. For retro-computing enthusiasts, this kind of analysis is replicable by mounting or hex-dumping period ISO images and comparing directory structures against known engine asset formats. It's a useful case study in constraint-driven engineering from an era before broadband patches or optional downloadable content existed.

Hacker News · 462 ptsRunnable

Cursor launches Origin, GitHub alternative

AI coding startup Cursor just built its own place to store and manage your code.

Cursor, the company behind a popular AI-powered code editor, has launched a new product called Origin that competes directly with GitHub — the dominant platform where developers store, share, and collaborate on code. Instead of just helping you write code with AI, Cursor now wants to also host the repositories where that code lives, potentially integrating AI features more deeply into the whole workflow. This matters because it signals AI coding tools moving from being an add-on to trying to own the entire developer pipeline, from writing to storing to shipping code. It's part of a broader trend of AI-native tools challenging established developer infrastructure.

Technical view

Cursor has released Origin, a git hosting and collaboration platform positioned as a GitHub alternative, extending its footprint beyond its AI-assisted code editor into source control and repository management. This lets Cursor tie AI-driven features (code review, PR generation, agentic workflows) more tightly into the hosting layer rather than depending on GitHub's APIs and extension points. Developers evaluating it would want to check migration tooling, CI/CD integration, and whether it supports existing git workflows or introduces a proprietary layer on top.

Hacker News · 437 ptsConceptual

Beware Management Consultants

A pointed essay on why hiring management consultants can backfire badly.

This is an opinion piece arguing that companies should think twice before bringing in management consultants — the outside experts firms hire to advise on strategy, restructuring, or efficiency. The argument likely centers on how consultants can offer generic, cookie-cutter advice, extract high fees, and sometimes leave organizations worse off, disconnected, or destabilized once they leave. It matters because consulting is a massive, influential industry that shapes decisions at governments and corporations worldwide, yet its actual track record is often debated. It's a reminder to be skeptical of outside authority dressed up in slides and frameworks.

Technical view

The piece appears to be a critical essay on the management consulting industry, likely covering incentive misalignment (consultants are paid regardless of outcomes), information asymmetry, and the tendency toward templated frameworks over context-specific solutions. Readers building organizational decision-making processes might take away the importance of internal capability-building and outcome-based engagement terms rather than pure advisory retainers. No specific data or case studies are given here, so treat the core claims as argumentative rather than empirically demonstrated.

Hacker News · 416 ptsConceptual

AI-Generated GitHub Copilot “Autofix” Allowed Compromise of Snowflake's Jira

An AI 'auto-fix' suggestion from GitHub Copilot reportedly helped attackers break into Snowflake's own Jira.

GitHub Copilot has a feature called Autofix that uses AI to automatically suggest patches for security vulnerabilities in code. This story describes a case where one of those AI-generated fixes actually introduced or failed to close a security hole, which attackers then used to compromise Snowflake's internal Jira system (the software many companies use to track bugs and projects). It's a cautionary tale about trusting AI-written security patches without careful human review, since a tool meant to protect code can sometimes create new risks. It matters because AI coding assistants are increasingly trusted with sensitive, security-critical work.

Technical view

The report details an incident where an AI-generated remediation from GitHub Copilot Autofix reportedly failed to properly close a vulnerability (or introduced a new flaw), which was subsequently exploited to compromise Snowflake's Jira instance. This is a concrete example of the risk in automated vulnerability-patching pipelines: AI-suggested fixes can pass superficial review while missing edge cases a human security engineer would catch. Practitioners running similar autofix tooling should treat AI-generated patches as draft suggestions requiring the same review rigor as manual PRs, especially for auth, access-control, or input-validation code paths.

Hacker News · 398 ptsBuildable

Using the railway network as a flatbed scanner

Someone turned Britain's entire rail network into a giant image scanner.

This is a quirky hacker project where someone repurposed data or sensors from the railway system to function like a flatbed scanner — the kind of device that captures an image line by line as it moves across a page. Instead of scanning paper, they likely used the physical movement of trains (and associated tracking or camera data) across the rail network to build up an image, turning infrastructure meant for transportation into an unconventional data-capture tool. It's the kind of playful, offbeat engineering that shows how creative people can repurpose everyday systems for entirely unintended uses. It matters less for practical utility and more as a delightful demonstration of lateral thinking in hardware and data hacking.

Technical view

The project likely exploits positional or imaging data generated as trains move along fixed rail routes — akin to a flatbed scanner's sensor head sweeping at a known, controlled speed — to reconstruct an image line-by-line from real-world railway telemetry or camera feeds. This falls into the 'weird sensor as scanner' genre of hardware hacks, similar to using satellite passes or disk drive heads as improvised scanning mechanisms. A reader interested in replicating this would need access to railway tracking data or trackside imaging hardware, plus signal-processing code to stitch swept captures into a coherent image.

Hacker News · 359 ptsRunnable

GPT 5.6 Sol is the best "vision" model OpenAI ever released

OpenAI says its new GPT-5.6 Sol model sees and understands images better than any before it.

GPT-5.6 Sol is described as OpenAI's strongest 'vision' model yet, meaning it's especially good at understanding and reasoning about images — not just recognizing objects, but interpreting charts, diagrams, handwriting, or complex visual scenes and explaining what's happening in them. This matters because a lot of real-world information isn't text: think medical scans, screenshots, whiteboard sketches, or photos of broken equipment, and a model that truly 'sees' well can help with tasks across science, coding, and everyday troubleshooting. The claim highlights how much AI vision capabilities have improved, closing the gap between how machines and humans interpret visual information. It's a milestone in the ongoing push to make AI assistants multimodal rather than text-only.

Technical view

GPT-5.6 Sol is positioned as OpenAI's best-performing vision-language model to date, implying improved benchmark results on tasks like visual question answering, chart/document understanding, OCR, and fine-grained visual reasoning compared to prior GPT releases. Practitioners building multimodal applications (document processing, visual agents, accessibility tools) would evaluate it via the API for tasks requiring precise grounding in image content rather than superficial captioning. Without further benchmark specifics in the abstract, the concrete gains should be validated against standard vision-language eval suites before relying on the claim for production decisions.

Hacker News · 354 ptsBuildable

Fixing a bricked Framework laptop

A step-by-step account of reviving a Framework laptop that wouldn't turn on.

Framework makes laptops designed to be modular and repairable, letting owners swap out parts instead of throwing the whole machine away — but this post is about a case where one still ended up 'bricked,' meaning it stopped working entirely, likely due to a firmware or hardware failure. The writer walks through how they diagnosed the problem and brought the laptop back to life, which is exactly the kind of repair story the right-to-repair movement champions. It matters because it shows both the promise and the limits of repairable hardware: even modular devices can fail in ways that require real troubleshooting skill. It's a practical, hands-on story for anyone who owns fixable tech or cares about reducing e-waste.

Technical view

This is a repair walkthrough documenting the diagnosis and recovery of a bricked Framework laptop, likely involving firmware/BIOS recovery, EC (embedded controller) reflashing, or board-level troubleshooting given Framework's modular mainboard design. Readers with similar hardware could replicate the fix by following the same diagnostic sequence — checking power delivery, reseating or reflashing firmware via recovery headers, and consulting Framework's community repair documentation. It's a useful reference for anyone doing low-level laptop firmware recovery work, particularly on modular/repairable hardware platforms.

Hacker News · 350 ptsRunnable

Fairphone is now officially available in the United States

Fairphone, the ethical repairable phone maker, finally sells officially in the US.

Fairphone is a company known for making smartphones designed to last longer and be easily repaired — you can swap the battery or screen yourself — while also trying to source materials more ethically than typical phone manufacturers. Until now, Fairphone has mainly sold in Europe, so US customers had to import devices unofficially; this news means Americans can now buy one through official channels, with proper warranty and support. It matters because it gives US consumers a real alternative to disposable flagship phones from Apple and Samsung, especially for people who care about sustainability and repairability. It's a small but meaningful expansion of the right-to-repair movement into the US market.

Technical view

Fairphone has launched official US retail availability, ending the prior situation where American buyers relied on gray-market imports or European resellers lacking local warranty and carrier support. This likely brings the device in line with US regulatory requirements (FCC certification, carrier band compatibility) and opens up manufacturer-backed repair parts and support domestically. For readers tracking modular/repairable hardware, this is a concrete data point in Fairphone's market expansion strategy and a test of US demand for repair-first smartphone design outside niche import channels.

Hacker News · 285 ptsConceptual

Field measurements of neighborhood-scale air temperature impacts of data centers

Data centers quietly warm up the neighborhoods around them, and now someone measured exactly how much.

Data centers packed with servers generate enormous heat, and much of it escapes into the surrounding air through cooling exhaust systems. Researchers wanted to know whether this heat noticeably warms the streets and homes nearby, not just inside the facility itself. They carried sensors through neighborhoods near data centers, taking real air-temperature readings at different distances and comparing them to areas farther away. The results show measurable local warming, adding a new environmental concern to the ongoing data center construction boom. This matters because as AI drives a surge in new data centers, nearby residents may face a hidden extra dose of urban heat on top of noise and power demands.

Technical view

Using field sensors (likely mobile transects or fixed loggers), researchers logged ambient air temperature at varying distances and directions from operating data center campuses to isolate their thermal footprint from broader urban heat-island effects. The approach controls for confounds like traffic and land cover by comparing near-facility and far-facility readings under similar weather conditions. Findings quantify a localized temperature increment attributable to data center waste-heat exhaust, likely from cooling towers or air-cooled chiller plants. Urban climate or siting practitioners could replicate the transect method to audit other point-source heat emitters, feeding results into zoning or cooling-system emission standards.

Hacker News · 278 ptsBuildable

Fairphone 6 and PostmarketOS working main camera

Getting a modern phone's camera to actually work on fully open-source Linux, driver by driver.

PostmarketOS is a project that ports mainline Linux, the same kind of operating system that runs servers and laptops, onto smartphones instead of their usual Android software. Phone cameras are notoriously hard to support this way because they rely on proprietary chips and secret vendor code that manufacturers rarely document. This report describes getting the Fairphone 6's main camera fully working under postmarketOS, likely by writing or adapting drivers through Linux's open camera framework. It matters because it pushes phones a step closer to being fully repairable, user-controlled computers rather than locked-down devices, echoing Fairphone's own mission of sustainable, modular hardware.

Technical view

Camera pipelines on modern smartphone SoCs typically need an image signal processor (ISP) driver, a sensor driver, and a userspace stack such as libcamera to expose a usable capture interface, all normally vendor-proprietary and tied to Android's HAL. Getting the Fairphone 6's main sensor working under postmarketOS implies maintaining out-of-tree (or newly upstreamed) kernel drivers for the sensor and ISP, plus a libcamera pipeline handler translating hardware-specific tuning into a standard API. This is meaningful progress for mainline Linux phone support, since camera subsystems are historically one of the last blockers to daily-driver usability. Developers can follow the project's device-tree and pipeline-handler patches to replicate the bring-up on similar SoCs.

Hacker News · 273 ptsConceptual

Linear algebra done right

The textbook that teaches linear algebra by thinking, not by grinding through determinants.

Linear algebra is the branch of math behind everything from computer graphics to machine learning, and it's usually taught by memorizing matrix procedures like determinants and row reduction. "Linear Algebra Done Right" is a well-known textbook that flips this approach, building the subject around the core ideas of vector spaces and linear maps, the abstract objects and transformations that matrices are just one way of representing. It leans on understanding why the rules work, using proofs and conceptual reasoning, rather than memorizing procedures for crunching numbers. It matters because this conceptual grounding tends to pay off later, in advanced math, physics, or machine learning theory, where thinking in terms of spaces and transformations, not grids of numbers, is what actually helps.

Technical view

The book (by Sheldon Axler) delays determinants until late in the text, instead developing the theory of finite-dimensional vector spaces, linear maps, eigenvalues, and inner product spaces directly, using invariant-subspace and minimal-polynomial arguments rather than determinant-based proofs for results like eigenvalue existence over complex fields. This departs from the standard computational curriculum and is popular as a rigorous second course for students who've already seen matrix mechanics. Anyone building intuition for topics like PCA, SVD, or the spectral theorem in ML or quant contexts can use it to shore up proof-level understanding that a purely computational course skips. It pairs well with problem sets or a rigorous course for self-study.

Hacker News · 266 ptsConceptual

Babies born under sugar rationing grew into adults with lower cancer risk

A wartime sugar shortage accidentally ran a decades-long experiment on cancer risk.

During and just after World War II, the UK strictly rationed sugar, meaning an entire generation of babies ate far less sugary food than the babies born right after rationing ended in 1953. Researchers realized this sudden, sharp before-and-after split is essentially a real-world experiment they could never have ethically designed on purpose. By comparing long-term health outcomes of people born just before versus just after rationing ended, they could isolate the effect of early-life sugar exposure from all the other lifestyle factors that usually make this kind of question impossible to study cleanly. The finding: the low-sugar babies grew up with a lower risk of developing cancer as adults. This matters because it's rare, strong evidence that what infants eat in their very first months may shape disease risk decades later.

Technical view

The study exploits the sharp 1953 discontinuity in UK sugar rationing as a natural experiment, comparing adult cancer incidence between cohorts conceived or born just before versus just after rationing ended, which produced an abrupt change in early-life sugar intake independent of confounds like parental income that changed more gradually. This regression-discontinuity-style design lets researchers attribute the cancer risk difference specifically to early-life sugar exposure rather than general nutrition or socioeconomic trends. The reported result is measurably lower cancer risk persisting into adulthood in the low-sugar-exposed cohort, adding to prior findings from this same cohort on diabetes and hypertension risk. Epidemiologists could extend this design to other rationing or policy-discontinuity events to further probe developmental-origins-of-disease hypotheses.

Hacker News · 262 ptsConceptual

Being ambitious and being a dad

Can you chase big career dreams and still make it home for bedtime? One dad's take.

This is a personal essay reflecting on the tension between having big ambitions, whether in a career, a startup, or some creative pursuit, and being present as a father. It grapples with a problem many working parents quietly face: time, energy, and attention are finite, and both ambition and parenting demand a lot of both. Rather than offering a formula, the piece works through the author's own experience to think about what it means to pursue big goals without sacrificing the relationship with your kids. It matters because it's a relatable, human meditation on a tradeoff many people struggle with but rarely discuss openly.

Technical view

As a first-person essay rather than a study, this doesn't present data or a method to replicate, but it works through a specific mental model: reframing ambition and fatherhood not as competing, zero-sum claims on time but as investments with different time horizons and feedback loops. Readers should expect qualitative reflections, such as concrete boundary-setting choices, rather than empirical claims. It's most useful as a prompt to examine one's own trade-offs and compare against similar essays in the founder-and-parent genre, with no external system or dataset to build on beyond adopting the author's framing.

Hacker News · 256 ptsConceptual

Claude Code May–August 2026 weekly limits promotion

Anthropic is loosening Claude Code's usage caps for a few months — here's the deal.

Claude Code is Anthropic's AI coding assistant, and like many subscription AI products it normally limits how much you can use it each week to control costs and server load. This item refers to a promotional stretch, running from May through August 2026, during which those weekly limits are adjusted, likely made more generous, to let people use the tool more freely. It's the kind of practical, timely update that mainly matters to people who already use or are considering Claude Code for programming work, since it directly affects how much coding help they can get before hitting a cap. It's less a scientific finding than an operational announcement about product limits.

Technical view

This references an Anthropic update to Claude Code's rate-limiting policy, specifically the rolling weekly usage caps governing how many tokens or requests a subscriber can consume before throttling, adjusted for the May–August 2026 window. Such promotions typically aim to drive adoption or gather usage data ahead of a pricing or limit change; practitioners should check their account's plan page or the Claude Code changelog for exact quota numbers, since specifics aren't given in the abstract. Teams relying on Claude Code for CI or heavy agentic workflows should treat this as a signal to reassess usage patterns before the promotional window ends, in case caps tighten afterward.

Hacker News · 249 ptsConceptual

On AI regulation and messaging

Anthropic's CEO weighs in on how the case for AI regulation should be made, not just written.

This links to a social media post by Dario Amodei, CEO of Anthropic (the company behind Claude), addressing AI regulation and, specifically, how the case for it should be communicated. The real problem here is that AI policy debates are often muddled: companies, politicians, and the public talk past each other about risk, innovation, and control, and how an argument is framed can matter as much as its substance. Rather than a technical proposal, this is Amodei's perspective on the rhetoric and strategy around pushing for sensible AI rules without stalling beneficial progress. It matters because Amodei is one of the most influential voices shaping how AI safety and regulation get discussed in Washington and beyond, so his framing choices can shift real policy.

Technical view

This is a link to a short-form social post (mirrored via xcancel, a Twitter/X front-end) from Anthropic CEO Dario Amodei on AI regulation and its public messaging, rather than a paper or artifact with a method to evaluate. Without the post's actual text, its concrete claims can't be summarized here, but it sits within Amodei's broader public stance advocating targeted AI safety regulation, balanced against concerns about regulatory capture or stifling smaller labs, echoing themes from his past essays and Anthropic's responsible scaling policy. Readers tracking AI policy should read the source post directly for the specific framing and any proposed messaging strategy.

Hacker News · 244 ptsConceptual

Meta Files Patent for Facial Recognition, Automatic Recording of People

A patent shows Meta glasses that could recognize your face and start filming automatically.

Meta, the company behind Facebook and Instagram, has filed a patent describing technology that pairs facial recognition with automatically starting a recording, likely aimed at its smart glasses like the Ray-Ban Meta line. The idea seems to be letting the device recognize a known face and then kick off video or audio capture without the wearer manually pressing a button. Patents describe possible future features, not shipped products, so this is a glimpse at what Meta's engineers are exploring rather than something available today. It matters because it raises immediate privacy questions: always-on smart glasses that can identify people by face and silently start recording would be a significant escalation in the surveillance capability tucked into ordinary-looking eyewear.

Technical view

The filing describes a system pairing facial recognition, on-device or cloud-based, with automatic triggering of audio/video capture, presumably integrated into Meta's smart glasses hardware, though patents typically claim broad functionality beyond what ever ships. Technically this would require a lightweight face-matching model, likely comparing against a user-curated contact list rather than open-world identification for latency and privacy-compliance reasons, plus a capture-triggering pipeline gated by recognition confidence. The significance lies less in novel technique and more in the product/legal signal: Meta actively pursuing IP around automated, identity-triggered recording, which will likely draw scrutiny under biometric privacy laws like Illinois' BIPA. Privacy researchers and policymakers should treat this as an early signal to monitor for eventual feature rollout, not evidence of a current shipping capability.

Hacker News · 230 ptsBuildable

Teaching my kid to code with a modern MUD

A parent turns an old-school text dungeon into a coding classroom for their kid.

A MUD is a 'multi-user dungeon' — a text-only online game from the 1980s where you type commands like 'go north' or 'attack dragon' to explore a shared world. This piece describes using a modern version of that format to teach a child programming, since MUDs are basically tiny worlds built entirely out of code that you can read, tweak, and extend. Instead of abstract exercises, the kid can add a new room, item, or monster and immediately see the result by playing the game. It matters because it turns programming from a dry syllabus into something with instant, playful feedback — a trick that's kept generations of programmers hooked.

Technical view

The author documents a hands-on curriculum using a modern MUD codebase (likely LPC/Ranvier/Evennia-style architecture) as a teaching vehicle for a child learning to program. Learners extend the game's object model — rooms, items, NPCs, command verbs — giving them a bounded but Turing-complete sandbox with immediate, observable output. This mirrors classic 'programming by tinkering' pedagogy (cf. Logo, Scratch) but with a text-adventure domain that naturally introduces state machines, event handling, and simple parsers. Replicable by picking an open-source MUD engine and scaffolding lessons around incremental feature additions.

Hacker News · 225 ptsConceptual

Ask HN: GitHub employees what's going on? Why?

Frustrated users ask GitHub staff to explain the site's repeated outages, off the record.

This is a forum post where someone asks current or former GitHub employees to explain why the site keeps having outages and problems, since outsiders only see press releases and guesswork. The poster is tired of speculation and wants real insight from people who actually work behind the scenes — what's breaking, why, and whether it's being fixed. It's less a research finding and more a community asking a big tech platform's own staff for honesty about its reliability. It matters to anyone who depends on GitHub daily, since unexplained outages affect software development worldwide.

Technical view

This is a discussion thread (Ask HN) soliciting first-hand accounts from GitHub engineers or staff about the root causes behind a pattern of recent reliability incidents, explicitly seeking insider technical detail over public postmortems or PR statements. No specific incident data or root-cause information is given in the prompt itself — the value is in whatever insider responses surface in the thread, which could range from infrastructure scaling issues to organizational factors post-acquisition. Readers interested in SRE practices would want to follow the actual comment thread for any substantive answers.

Hacker News · 224 ptsConceptual

Rethinking Database Programming

A fresh take on how programmers should talk to databases from their code.

Databases and programming languages have historically felt like two separate worlds glued together with clunky query strings (like SQL embedded inside code) — this piece argues for rethinking that relationship. The core problem is that writing database logic often means switching mental modes, juggling type mismatches, and losing the safety nets your programming language normally gives you. The proposed approach likely involves treating database operations as first-class, type-checked parts of the language itself rather than bolted-on strings, so the compiler can catch mistakes early. It matters because most software bugs and security holes cluster right at this database boundary, so smoothing it out has broad payoff for reliability.

Technical view

The piece proposes reconsidering the interface between application code and databases, likely arguing against the traditional string-based SQL-embedded-in-host-language model in favor of tighter integration — such as language-integrated query, type-safe query builders, or algebraic effect-style abstractions over persistence. This connects to a lineage of work like LINQ, Rust's Diesel/SeaORM, and research on gradual/staged query compilation, where the goal is compile-time verification of queries against schema and stronger composability. A practitioner could evaluate the argument against their own ORM/query-builder stack and consider adopting type-checked query DSLs to catch schema mismatches before runtime.

Hacker News · 222 ptsConceptual

How does IKEA come up with names for its products?

Ever wonder why IKEA furniture has names like Billy or Kallax? Here's the system.

IKEA famously names its products things like 'Billy' bookcase or 'Kallax' shelving instead of model numbers, and this piece explains the method behind it. The furniture giant actually has an internal naming system where different categories of products draw from different theme lists — Swedish place names, first names, or words related to nature, depending on what the item is. The approach exists partly out of necessity: founder Ingvar Kamprad was dyslexic and found names easier to remember than codes, and it also gives the massive catalog a distinctly Scandinavian, friendly identity worldwide. It matters as a neat case study in branding and information design at a global scale — tens of thousands of products need a memorable, consistent way to be told apart.

Technical view

The article details IKEA's internal product-naming taxonomy, in which categories map to specific naming conventions (e.g., bookcases get Swedish boys' names, fabrics get women's names, outdoor furniture gets Scandinavian island names), a system reportedly originating from founder Ingvar Kamprad's dyslexia making numeric SKUs hard to track. This is essentially a controlled-vocabulary classification scheme applied at retail scale, functioning as both a mnemonic device and a brand-consistency mechanism across tens of thousands of SKUs and dozens of languages. Readers interested in taxonomy or naming systems for large catalogs can draw parallels to controlled vocabularies used in software versioning or product-line branding.

Hacker News · 206 ptsConceptual

Norway should buy OpenAI

An op-ed argues Norway's oil fortune should go toward literally owning OpenAI.

Norway has one of the world's largest sovereign wealth funds, built from decades of oil profits, and this opinion piece argues that fund should be used to buy a stake in — or outright acquire — OpenAI, the company behind ChatGPT. The reasoning is likely that AI is becoming as strategically important as oil once was, so a nation with capital to spare should secure ownership in the technology shaping the future rather than just consuming it. This is a bold, hypothetical proposal rather than an announced deal — more a thought experiment about national strategy in the AI race. It matters because it raises real questions about who should control transformative AI: private Silicon Valley investors, or sovereign nations acting on behalf of their citizens.

Technical view

This is an opinion piece proposing that Norway's Government Pension Fund Global (the ~$1.7 trillion sovereign wealth fund) take a strategic ownership position in OpenAI, framing AI compute and model ownership as a resource comparable to the oil wealth that built the fund. No concrete deal mechanics, valuation, or governance structure are specified in the abstract — the argument is at the level of national industrial strategy rather than a transaction proposal. It sits within a broader debate about sovereign AI investment (cf. UAE's MGX, Saudi PIF stakes in AI infrastructure) and could be read alongside analyses of state actors hedging against concentrated private control of frontier AI labs.

Hacker News · 202 ptsBuildable

Turbovec – Google's TurboQuant for vector search in Rust

A Rust rewrite of Google's clever trick for searching billions of vectors fast and cheap.

When apps do things like 'find similar images' or power AI search, they compare huge lists of numbers called vectors, and doing this fast at massive scale is a hard engineering problem. Google published a technique called TurboQuant that compresses these vectors cleverly so you can search through billions of them without needing huge amounts of expensive memory. Turbovec is a project that reimplements that same idea in Rust, a programming language prized for speed and safety, making the technique usable outside Google's internal systems. It matters because fast, affordable vector search underpins a lot of modern AI features — recommendation engines, semantic search, chatbots retrieving relevant facts — and open implementations let more developers use these tricks.

Technical view

Turbovec is an open-source Rust reimplementation of Google's TurboQuant vector quantization technique, aimed at reducing the memory footprint of approximate nearest-neighbor (ANN) search over large embedding datasets while preserving search quality. Quantization approaches like this typically compress float vectors into lower-bit representations (e.g., product quantization or scalar quantization variants) to shrink index size and improve cache locality, trading a small accuracy loss for large speed/memory gains at billion-scale. A practitioner building a vector database or RAG pipeline could integrate Turbovec as a compression layer ahead of an ANN index (HNSW, IVF) to cut infrastructure costs, and compare its Rust performance characteristics against existing quantization libraries like Faiss.

Hacker News · 177 ptsConceptual

India has paved the way for charging merchants a fee on UPI transactions

India just let banks start charging shops for using its wildly popular free payment app.

UPI (Unified Payments Interface) is India's massively popular instant payment system that lets people pay for anything by scanning a QR code, and until now it's been free for merchants to accept. This news is about India's regulators opening the door for banks or payment providers to start charging merchants a fee on those transactions, a shift from the free model that helped UPI become so widely adopted. The real-world tension is between keeping digital payments accessible for small shopkeepers versus letting the companies running the infrastructure actually make money to sustain it. It matters because UPI processes billions of transactions and is often cited as a model for other countries — how India balances 'free for growth' versus 'sustainable business' will influence payment systems globally.

Technical view

India's regulators/payment council have approved allowing UPI transaction fees to be charged to merchants, reversing or modifying the zero-Merchant Discount Rate (MDR) policy that has applied to UPI person-to-merchant payments since 2020. This affects the economics of the National Payments Corporation of India (NPCI)-run UPI rail, which processes tens of billions of transactions monthly, and has implications for how banks, PSPs (payment service providers like PhonePe/Google Pay), and fintechs recoup infrastructure costs. Practitioners in payments/fintech should watch for the specific fee structure and thresholds (e.g., exemptions for small merchants) as this shapes competitive dynamics between UPI and card networks in India.

Hacker News · 172 ptsConceptual

Repair Cafe – Fix Your Broken Items

Volunteer-run cafes where strangers help you fix your broken stuff instead of tossing it.

A Repair Cafe is a community event, often held regularly in a local space, where volunteers with fix-it skills help people repair broken household items — toasters, lamps, clothes, bikes — for free instead of throwing them away. The problem it addresses is the throwaway culture around cheap electronics and goods, where it's often easier to buy new than to fix old, generating waste and losing repair knowledge. The approach is simple and social: bring your broken item, sit with a volunteer fixer, and learn how it's done together rather than just dropping it off. It matters because it saves money, reduces landfill waste, and keeps practical repair skills alive in an era when many products are designed to be replaced rather than mended.

Technical view

Repair Cafes are grassroots, volunteer-staffed community events (originating from a Dutch initiative and now an international network) where attendees bring broken consumer items and work alongside skilled volunteers to diagnose and fix them on-site, emphasizing skill transfer over drop-off service. This model addresses e-waste reduction and 'right to repair' concerns by countering planned obsolescence and lack of consumer repair documentation, functioning as informal knowledge-sharing infrastructure outside manufacturer service channels. Anyone can replicate the model by organizing a recurring local meetup with tools, multimeters, and volunteers experienced in electronics, textiles, or small appliance repair, with several existing directories (e.g., Repair Café International Foundation) offering event templates and locator maps.

Hacker News · 170 ptsConceptual

The Benchmarkpocalypse

AI benchmarks are quietly collapsing as models learn to game the test.

For years, AI progress was tracked by scores on standardized tests — solve more math problems, answer more trivia, and you're 'better.' But labs increasingly train on data that overlaps with these tests, or optimize specifically for the score rather than real ability, so a high number stops meaning what it used to. This piece is about that unraveling: benchmarks that were supposed to be objective yardsticks turning into a race everyone can win without actually improving. It matters because businesses, researchers, and users all lean on these scores to pick models, and if the scores lie, so does everyone's sense of what's actually good.

Technical view

The core issue is Goodhart's law applied to LLM evaluation: once a benchmark becomes a target, it stops being a reliable measure, whether through direct contamination (test data leaking into training corpora), overfitting to a benchmark's specific quirks, or score saturation where every frontier model clusters near ceiling. Practitioners increasingly compensate with held-out or continuously refreshed private eval sets, human-preference arenas (à la LMSYS/Chatbot Arena), and task-specific real-world evals rather than static leaderboards. Anyone building eval infrastructure should assume public benchmark numbers are a weak signal and budget for custom, contamination-resistant test suites.

Hacker News · 162 ptsConceptual

Los Puesteros, solitary men who look after ranches and livestock in Patagonia

Meet the last solitary herders who spend months alone guarding sheep in windswept Patagonia.

A 'puestero' is a caretaker who lives alone, often for months at a stretch, on a remote outpost of a Patagonian ranch (estancia), watching over sheep and cattle across vast, empty land. The piece follows their day-to-day existence — isolation, harsh weather, and a self-sufficient way of life that's largely unchanged for generations — as a way of documenting a disappearing rural tradition. It approaches this through storytelling and observation rather than data, letting the puesteros' routines and solitude speak for themselves. It matters as a window into a vanishing lifestyle being squeezed out by depopulation, changing land economics, and the pull of cities.

Technical view

This is a documentary/human-interest piece rather than a technical one; its substance lies in ethnographic observation of a specific labor practice tied to extensive Patagonian ranching. The implicit claim is that this way of life is under pressure and worth recording before it fades. A reader wanting to build on it would look toward oral history archiving, rural economic studies of Patagonian sheep farming, or comparative documentation of solitary pastoral labor elsewhere in the world.

Hacker News · 161 ptsRunnable

Show HN: Desktopcolors.com – A museum for solid background colors of classic OS

A tiny web museum preserves the flat, iconic desktop background colors of old operating systems.

The creator built a small website that collects the plain solid-color backgrounds classic operating systems used before fancy wallpapers existed — think the specific teal of old Windows or the gray of early Mac OS. It's a nostalgia project made during a vacation, essentially a gallery where each color is labeled by which OS and version it came from. There's no deep technical approach here; it's a simple, lovingly curated reference site. It matters as a fun bit of digital preservation, capturing tiny design details that shaped how millions of people experienced their computers.

Technical view

Implementation-wise this is almost certainly a static site pairing each OS/version with its exact background hex or RGB value, likely sourced from old OS documentation, screenshots, or source assets. It's a low-complexity personal project, but a nice pattern for design-history archiving; one could extend it by adding resolution/year metadata, letting users submit missing colors, or generating a downloadable palette/swatch file for retro-UI projects.

Hacker News · 159 ptsBuildable

Python Polars Cheatsheet (based on our O'Reilly book)

A cheat sheet for Polars, the Rust-powered dataframe library that's leaving pandas in the dust.

Polars is a library for working with spreadsheet-like data tables in Python, similar to the popular pandas library but built in Rust for speed, using all your computer's CPU cores at once instead of just one. This cheat sheet, drawn from an O'Reilly book on the subject, condenses the common things you'd want to do — filtering rows, grouping and summarizing, joining tables — into a quick reference. The approach is simply distillation: take a full book's worth of technique and compress it to the essentials you'd want on hand while coding. It matters because as datasets grow, Polars' speed advantage makes it an increasingly common upgrade path for data analysts hitting pandas' limits.

Technical view

Polars stores data in Apache Arrow's columnar memory format and offers a lazy evaluation API, letting it build and optimize a full query plan (predicate pushdown, projection pushdown) before executing, rather than running operations eagerly line by line like pandas. The cheatsheet packages its expression-based syntax for select/filter/group_by/join operations as quick-lookup reference. A practitioner could use it as a bridge document while porting pandas pipelines to Polars, particularly for CPU-bound ETL jobs where multithreading yields large speedups.

Hacker News · 157 ptsConceptual

And then the men with guns tell you to do it anyway

When the law shows up armed, your objections stop mattering — you comply anyway.

The title alone signals a story about being forced to do something against your will or judgment once armed authority gets involved — the kind of moment where principle runs into raw state power and loses. Without more detail, the specific context (a legal order, a government demand on a company, a personal encounter with police or military) isn't given, but the theme is the gap between what you think is right and what you're compelled to do when someone with guns is the one asking. It approaches this through narrative rather than analysis, telling a specific incident rather than arguing an abstract point. It matters because it captures something people rarely admit outright: resistance has limits, and coercive power tends to win in the moment even when it's wrong.

Technical view

No abstract is available, so specifics of the mechanism or claim can't be confirmed — this reads as a first-person or narrative account of coerced compliance under state or armed authority, a theme that recurs in discussions of surveillance mandates, forced disclosure, or law-enforcement encounters. Readers interested in the underlying issue (e.g., legal compulsion of individuals or companies) would need to read the full piece to determine which specific situation it describes before drawing technical or policy conclusions.

Hacker News · 153 ptsBuildable

Claude writing a macOS driver for my obscure HP printer built only for Windows

An AI wrote a working Mac driver for a printer whose maker only ever supported Windows.

The author owned an old HP printer that HP never bothered to support on macOS, leaving it effectively useless on a Mac. Instead of giving up, they used Claude, an AI coding assistant, to figure out how the printer communicates and write the missing driver software from scratch. The AI worked through the unfamiliar printer's command format and built code that lets macOS talk to it properly, essentially reverse-engineering hardware documentation that never existed. It matters as a vivid example of AI coding tools tackling messy, real-world problems — not just writing boilerplate, but solving a niche technical puzzle no company found profitable enough to fix.

Technical view

This likely involved inspecting the printer's USB or network communication (packet captures, PCL/PostScript-like command sequences) and having Claude synthesize a macOS-compatible driver, probably as a CUPS (Common Unix Printing System) filter or PPD-equivalent, since macOS printing is built on CUPS. It's a concrete demonstration of LLM-assisted reverse engineering of an undocumented protocol producing functional low-level system software. Anyone facing a similar orphaned-hardware problem could replicate the approach: capture the print job's raw byte stream, feed the patterns to an AI assistant, and iterate toward a working CUPS driver.

Hacker News · 151 ptsRunnable

Finger: Social network that never died

A 1970s one-line Unix command for checking who's online never actually went away.

Finger is an ancient internet protocol, older than the web, that let you type a simple command to see who else was logged into a shared computer and read their self-written status update (a '.plan' file) — essentially a text-only status feed decades before Facebook or Twitter existed. The piece explores how this tiny, stripped-down tool quietly survived as a niche 'social network' among hobbyists and retro-computing fans who still run finger servers today. Its approach is minimalism itself: no likes, no algorithm, just a name and a short plain-text message. It matters as a reminder that the core idea of 'social status updates' predates modern social media by decades, and that simplicity has its own staying power.

Technical view

Finger is defined by RFC 1288 and traditionally runs over TCP port 79, returning a user's login status plus the contents of their `.plan`/`.project` files from a Unix account. Small self-hosted finger servers persist today among retro-computing and hacker-culture communities as an intentionally minimal alternative to modern social platforms. A reader could stand up their own finger daemon (e.g., `fingerd` or modern reimplementations) in an afternoon and get a working, decades-old social protocol running on a home server.

Hacker News · 146 ptsConceptual

Degraded performance for multiple models

A status-page alert: several AI models are running slower or glitchier than normal right now.

This is essentially an incident notice — the kind of thing a cloud provider posts when something behind the scenes isn't working smoothly — reporting that multiple AI models are experiencing 'degraded performance,' meaning responses might be slower, more likely to fail, or otherwise less reliable than usual. These issues typically stem from strained computing capacity, infrastructure hiccups, or unexpected demand spikes on the servers running the models. There's no deep technique to explain here, just a transparency notice letting users know something's off so they can adjust expectations. It matters practically: anyone building apps on top of these models may see slowdowns or errors and should know it's a known, provider-side issue rather than a bug in their own code.

Technical view

Status-page 'degraded performance' incidents for hosted LLMs typically reflect elevated latency, increased error/timeout rates, or partial capacity loss in the inference serving infrastructure, sometimes cascading across multiple models sharing backend resources. Standard practitioner mitigations include implementing retry-with-backoff logic, routing across model or provider fallbacks, and monitoring the provider's status page/API for recovery signals before resuming full traffic. This kind of notice is operational rather than technical content — its main value is confirming the issue is upstream, not in your integration.

Hacker News · 143 ptsConceptual

A particle made of force: physicists say they've found mysterious 'glueball'

Physicists say they've caught a particle made purely of the force that binds atoms.

Deep inside atoms, gluons are the messenger particles that glue quarks together into protons and neutrons via the 'strong force.' A glueball is a particle built entirely out of gluons sticking to each other, with no quarks at all — a strange state predicted decades ago by the theory of the strong force but never confirmed. Physicists hunt for it by smashing particles together and sifting the debris for something with the right mass and behavior but no quark fingerprint. Nailing one down would be direct proof that force-carriers alone can clump into matter, a genuinely weird idea made real.

Technical view

Glueballs are bound states of gluons predicted by QCD (quantum chromodynamics), expected in the 1.5–2.5 GeV mass range but notoriously hard to isolate because they mix with ordinary quark-antiquark mesons of the same quantum numbers. Candidate signals like f0(1500) and f0(1710) have circulated for years; new claims typically come from precision spectroscopy in collider or fixed-target data (e.g., BESIII, LHCb) via partial-wave analysis of decay channels. Practitioners can dig into the branching ratios and production mechanisms cited to judge how cleanly quark-model backgrounds were excluded.

Hacker News · 141 ptsRunnable

A 3D fruit fly on macOS desktop powered by the real FlyWire connectome

A virtual fruit fly roams your Mac desktop, its brain wired exactly like a real one's.

FlyWire is a massive research project that mapped every neuron and connection in a real fruit fly's brain — its 'connectome.' This project takes that actual wiring map and uses it to animate a 3D fly that lives on your desktop, so its movements come from simulating real neural circuitry rather than an artist's hand-coded animation. It's a playful way to make cutting-edge neuroscience visible and interactive. It also hints at a future where digital creatures are built from real biological blueprints instead of guesswork.

Technical view

The FlyWire consortium reconstructed the full synaptic connectome of the Drosophila brain (~140,000 neurons) from electron-microscopy imaging; this desktop app appears to drive a 3D avatar's behavior from that connectivity data rather than scripted motion. Builders interested in computational neuroscience could inspect how neural activity is simulated and mapped to movement, and potentially swap in modified circuits to test hypotheses about specific neurons' roles.

Hacker News · 140 ptsConceptual

California's new tire efficiency rules could save drivers $1B a year

California's new tire rules could quietly save drivers a billion dollars a year in gas.

Tires that resist rolling less make a car easier to push down the road, which means less fuel burned or more range for an EV — like the difference between pushing a bike with soft tires versus firm ones. California is setting new minimum efficiency standards that tire makers must meet to sell replacement tires in the state, nudging the whole market toward lower-friction designs. Because every driver eventually buys new tires, the savings add up fast across millions of cars. State regulators estimate the combined fuel savings at roughly $1 billion annually.

Technical view

The rules set minimum rolling-resistance-coefficient thresholds for replacement tires sold in California, likely enforced by state regulators as an extension of existing efficiency/low-carbon transportation policy. Tire manufacturers will need to adjust compounds and tread design to meet compliance, which could shift national product lines given California's market size. The $1B/year figure is presumably derived from aggregate fuel-cost reduction across the state's vehicle fleet.

Hacker News · 138 ptsConceptual

US announces new sanctions on top ICC figures

Washington slaps sanctions on senior officials of the world's war-crimes court.

The International Criminal Court (ICC) is a global body that investigates and prosecutes war crimes and crimes against humanity. The US government has imposed sanctions — things like frozen assets and travel bans — on top ICC officials, continuing a dispute over the court's authority to investigate US allies. This escalates tension between Washington and an institution most of the world recognizes but the US has never fully joined. It matters because it affects how international justice gets enforced and strains relations with allied nations who support the court.

Technical view

The sanctions likely follow the pattern of prior executive actions targeting ICC prosecutors and judges over investigations or arrest warrants involving US or allied officials, using asset-freeze and entry-ban authorities. This continues a recurring US-ICC standoff rooted in disputes over jurisdiction over non-member states. Analysts would want to track which specific officials are named and the stated legal basis to gauge scope and likely diplomatic fallout.

Hacker News · 126 ptsBuildable

Mojo is now open source!

A fast, Python-friendly AI programming language just threw open its source code.

Mojo is a programming language built to feel as easy to write as Python but run as fast as low-level languages like C++, aimed especially at AI and high-performance software. Until now, its inner workings were kept private by its creator, Modular; now the code itself has been released publicly so anyone can read, modify, and contribute to it. This matters because it lets outside developers trust, improve, and extend the language directly rather than waiting on one company, potentially speeding up how fast useful features and hardware support arrive.

Technical view

Mojo is a Python-superset language built on MLIR, offering static typing, memory-safety features, and direct hardware targeting (CPU/GPU) for performance-critical AI workloads without leaving the Python ecosystem. Open-sourcing the compiler and runtime lets developers inspect codegen, contribute backend optimizations, and build custom kernels or hardware targets rather than relying solely on Modular's roadmap. Practitioners can now fork or upstream changes and audit how Mojo achieves near-native performance from Python-like syntax.

Hacker News · 120 ptsBuildable

Splitting a Git Commit

A guide to slicing one messy code change into several clean, reviewable commits.

When coding, a 'commit' is a saved snapshot of your changes; sometimes you bundle too much into one commit and later wish you'd split it into several logical pieces. This piece walks through how to carefully break a single commit into multiple smaller ones without losing any work. It relies on Git's built-in tools that let you selectively pick which lines of code go into which commit. This matters because a clean, well-organized history makes it far easier for others (and future you) to review, understand, and debug the code later.

Technical view

The technique centers on `git rebase -i` to mark a commit for `edit`, then `git reset HEAD^` to unstage its changes, followed by `git add -p` to interactively stage individual hunks into separate commits before running `rebase --continue`. This is standard practice for cleaning history pre-review or isolating changes for `git bisect`. Anyone doing code review workflows can apply this directly to turn a monolithic diff into a reviewable, logically-ordered commit sequence.

Hacker News · 117 ptsConceptual

How to put 170 atoms in an atom

Scientists pack 170 individual atoms into a single quantum computing chip using lasers.

Quantum computers built from 'neutral atoms' work by using tightly focused laser beams — called optical tweezers — to grab and pin individual atoms in place, arranging them into precise grids where each atom acts as a quantum bit, or qubit. The title's pun refers to loading 170 of these physical atoms into one such device. Doing this reliably is hard: atoms are constantly jostled and must be rearranged and cooled with exquisite laser control to sit exactly where needed. Packing in more atoms without losing precision is one of the main paths to building bigger, more powerful quantum computers.

Technical view

This describes scaling neutral-atom quantum computing hardware, where optical tweezer arrays trap individual atoms (often alkali or alkaline-earth species) and rearrange them via real-time feedback to fill defect-free lattice configurations. Reaching 170 atoms in a single array reflects incremental progress in tweezer count, trapping fidelity, and atom-loading/rearrangement algorithms — key bottlenecks for qubit-count scaling in this architecture. Readers tracking quantum hardware roadmaps can compare this atom count against competing platforms (superconducting, trapped-ion) to gauge relative scaling trajectories.

Hacker News · 114 ptsRunnable

Launch HN: Speko (YC S26) – OpenRouter for Voice AI

A startup finds the cheapest, best-sounding combo of AI models for your voice app.

Building a voice AI assistant usually means stringing together three separate AI models: one that turns your speech into text, one (a language model) that figures out a smart response, and one that turns that text back into natural-sounding speech. There are dozens of options for each piece, quality and prices shift monthly, and swapping any one out means painful re-integration, so most teams pick once and get stuck running outdated, pricier models. Speko benchmarks all these options together and automatically recommends the best combination for your specific needs and budget, explaining why. It's essentially a matchmaking layer that keeps a voice AI product on the best available tech without constant manual rework.

Technical view

Speko is a benchmarking and recommendation platform spanning the full voice-agent pipeline — STT, LLM, and TTS — that scores public vendor options against user-specified constraints (cost, latency, quality) and surfaces an optimal stack with justification. It positions itself analogously to OpenRouter's model-routing abstraction but extended across three chained model types instead of one, aiming to remove the integration friction that currently locks teams into stale vendor choices. Teams building voice agents could use it to continuously re-evaluate and swap components as better/cheaper models ship, without re-engineering the pipeline each time.

Hacker News · 112 ptsConceptual

Apple announces changes for apps in the European Union

Apple adjusts iPhone App Store rules in Europe to comply with EU tech law.

Apple is changing how its iPhone app store works, but only for people in the European Union, because a new EU law called the Digital Markets Act forces big tech companies to give users and developers more choice — like installing apps from outside Apple's own store or using other payment systems. Apple has resisted many of these changes, arguing they hurt security and its business model, so this announcement is another round in an ongoing back-and-forth over how much control Apple must give up. The 'how' here is regulatory, not technical: EU regulators set rules, and Apple responds with policy tweaks to fees, app review, and sideloading options. It matters because it could reshape how much power any single company gets over what software billions of people can install on their phones.

Technical view

Apple is updating App Store and platform policies specifically for iOS/iPadOS in the EU to comply with Digital Markets Act (DMA) obligations, likely touching alternative marketplaces, sideloading, browser engine choice, and the Core Technology Fee structure. Developers targeting EU users should watch for updated entitlements, notarization requirements for third-party app stores, and fee schedules that differ from Apple's global terms. This continues Apple's pattern of a bifurcated compliance posture (EU vs. rest-of-world) rather than a single global policy. Practitioners distributing iOS apps in Europe should review Apple's developer documentation for the specific new terms before the next enforcement deadline.

Hacker News · 110 ptsConceptual

Deus Ex creator Warren Spector is retiring from game development

The mind behind Deus Ex and Thief is stepping away from making games.

Warren Spector is one of the most influential video game designers of the last 30 years, best known for creating Deus Ex, a genre-defining game that let players solve problems through stealth, hacking, or combat instead of one fixed path, and for shaping immersive-sim games like Thief and System Shock. He's announced his retirement, closing out a career built around giving players real choices rather than a single scripted route through a story. There's no new technology here — it's a career milestone — but it matters because his design philosophy quietly influences huge swaths of modern game design, from Dishonored to Baldur's Gate 3. His departure marks an end-of-era moment for a specific, beloved school of game design thinking.

Technical view

Spector, lead designer/producer on Deus Ex (2000), Thief: The Dark Project, and System Shock, is retiring, closing out a multi-decade career central to the 'immersive sim' genre — games built around simulated systems, emergent interactions, and multiple viable solution paths rather than scripted linear design. The retirement introduces no new methods but caps a body of work that directly informs contemporary titles (Dishonored, Prey, Baldur's Gate 3) built on similar systemic-design principles. There's no mechanism to build on beyond his design writings and GDC talks as reference material for immersive-sim architecture.

Hacker News · 109 ptsBuildable

Claude Code Teaching macOS to Natively Print to the HP Laser 1008a

Someone used an AI coding agent to hack together printer support macOS never shipped.

This is a story about a person using Claude Code, an AI assistant that writes and runs code, to solve an annoying real-world problem: their HP LaserJet 1008a printer isn't natively supported by macOS, so instead of hunting for old drivers, they had the AI help build the missing software glue — likely a print driver file that plugs into macOS's underlying printing system — to make the printer work. This matters as a demonstration of AI coding tools tackling messy, unglamorous 'long-tail' compatibility problems that companies never bothered to fix, letting one individual patch a gap in an operating system's hardware support on their own.

Technical view

The write-up documents using Claude Code as an agentic pair-programmer to build native macOS print support for the HP LaserJet 1008a, a model lacking official Apple/HP driver support, likely involving CUPS filter or PPD creation and reverse-engineering the printer's communication protocol from documentation or packet captures. The interesting technical takeaway is the workflow: using an LLM agent to iteratively inspect printer communication, write driver code, and debug against real hardware in a loop, rather than the driver logic itself being novel. Anyone with an unsupported peripheral could replicate the approach by pointing an agentic coding tool at protocol docs/traces and iterating toward a working CUPS filter.

Hacker News · 103 ptsConceptual

The Road to MS-DOS 2.0

A deep dive into how MS-DOS grew up between its scrappy 1.0 and capable 2.0 releases.

This is a history piece tracing how Microsoft's original PC operating system, MS-DOS, evolved from its bare-bones first version into the far more capable 2.0 release, which added things like a hierarchical file system (folders within folders, obvious now but not yet standard then) and support for IBM's PC XT hardware. The 'how' is historical storytelling — pulling from old source code and design decisions to show why certain features were added and what constraints, like memory limits and hardware compatibility, shaped them. It matters to anyone curious about the ancestry of modern operating systems, since PC-era conventions from DOS still echo in Windows and command-line habits used today.

Technical view

The piece is a historical retrospective on the development of MS-DOS 2.0 (1983), which introduced a hierarchical directory structure borrowing ideas from Unix, device drivers, and support for the IBM PC XT's 10MB hard disk, moving well beyond DOS 1.0's flat file system and floppy-only design. Such retrospectives typically draw on Microsoft's since-open-sourced early MS-DOS code and design documents to reconstruct engineering trade-offs made under tight memory and hardware constraints. For practitioners, it's a useful case study in incremental OS design evolution and the API compatibility decisions that persisted for decades in DOS and Windows.

Hacker News · 99 ptsConceptual

Exercise intensity modulates interorgan communication and is associated with

How hard you exercise changes the chemical messages your organs send each other.

This study looks at how the body's organs 'talk' to each other during exercise — muscles, fat tissue, the liver, and others release signaling molecules like hormones or small proteins into the blood that tell other organs how to respond, a process called interorgan communication. The researchers find that exercise intensity, how hard you're working, not just how long, changes which of these signals get sent and how strongly, suggesting the body has different response 'modes' depending on effort level. They likely measured these molecules in blood samples taken during exercise at different intensities and linked the patterns to health-related associations. This matters because understanding these signals could help scientists design better exercise guidance or even drugs that mimic exercise's health benefits for people who can't exercise much.

Technical view

The study characterizes intensity-dependent interorgan crosstalk during exercise, profiling circulating signaling factors, likely myokines, adipokines, or hepatokines, across different exercise intensities to map how metabolic and signaling responses scale with effort. The approach probably involves controlled, graded-intensity exercise protocols paired with blood sampling and targeted or omics-based assays to quantify secreted factors, then statistically associating specific signaling patterns with physiological or health outcomes. This adds to the growing 'exercise as medicine' molecular literature, and practitioners in exercise physiology or metabolic disease research could use the identified intensity-signal associations to refine dosing of exercise interventions or pursue exercise-mimetic drug targets.

Hacker News · 87 ptsConceptual

Shattered skeleton is first confirmed death from trebuchet

Archaeologists say a shattered medieval skeleton is the first proven trebuchet kill.

Archaeologists examined a skeleton with severe, unusual bone damage and concluded it's the first confirmed case of someone killed by a trebuchet, the massive medieval siege weapon that flung heavy projectiles at castle walls and defenders. The 'how' of the discovery is forensic: researchers studied the pattern and severity of the fractures, which don't match typical sword, arrow, or blunt-weapon injuries, and matched them to the kind of massive blunt force a trebuchet projectile would cause, likely cross-referencing historical siege records to place the death at a known battle. It matters because it turns an object mostly known from documents and reconstructions into something with direct physical evidence of its lethality, adding real data to how historians understand medieval siege warfare.

Technical view

Researchers performed forensic skeletal analysis on human remains showing catastrophic, high-energy blunt trauma inconsistent with contemporary edged or projectile weapons, attributing the injury pattern to trebuchet projectile impact, likely the first such case confirmed via osteological/forensic methods rather than textual inference alone. The methodology presumably combines fracture pattern analysis, such as comminution and energy transfer signatures, with archaeological context like site dating and associated siege history to rule out alternative causes of death. This contributes a physical-evidence data point to medieval military archaeology and could inform biomechanical modeling of historical siege weapon lethality for further comparative studies.

Hacker News · 84 ptsBuildable

A 25-year-old video patent just expired, ending a legal headache for Linux

An old video-codec patent just died, quietly freeing Linux from a decades-long legal worry.

For 25 years, a patent covering a piece of video technology has loomed over open-source software, meaning developers had to tread carefully or leave out certain video features in Linux to avoid lawsuits — a common headache in the codec world, where key pieces of formats like MPEG have been patent-encumbered even though the underlying code is easy to write. That patent has now expired, and under patent law, protection lasts about 20 years from filing, so anyone, including free and open-source projects, can now implement that technique without fear of being sued. The 'how' isn't a technical fix, it's just the passage of time running out a legal clock. It matters because it removes a legal landmine that forced distros and developers into workarounds, license fees, or missing features for a quarter century.

Technical view

A patent covering a specific video encoding/decoding technique, likely tied to an MPEG or related codec component, has reached the end of its roughly 20-year enforceable term, removing a longstanding legal risk for Linux distributions and open-source multimedia projects that previously had to disable, patch around, or ship the feature only as an optional, non-default package. This is a legal rather than technical event, but it directly affects what maintainers can now ship enabled by default without patent licensing exposure, for example in ffmpeg, GStreamer, or distro media stacks. Practitioners maintaining codec support in open-source projects can now revisit previously patent-gated code paths and enable them by default, and should check whether related patents in the same family have also lapsed.