Deep-Tech Digest // 2026-08-10 · IST

Monday, 10 August 2026

168 new items across 5 fields — pulled from arXiv & Hacker News, deduped against everything served before. Each card gives you the problem, how it works, and what’s new in plain words — read that first; open Go deeper only when a card earns it.

46AI & Machine Learning
20Robotics
45Biology
1Quanta — Explained
56What's Trending
AI

AI & Machine Learning

46 new
arXiv · cs.CLBuildable★ flagship

Learning When to Trust via Selective Context Preference Optimization

Teaching AI when to believe outside hints and when to ignore them.

When you give an AI extra context — a retrieved document, a note, a hint — that context can help it or fool it. If the AI is trained to just distrust everything, it becomes safe but stubborn and stops using genuinely helpful information; if it trusts everything, one misleading line can flip a right answer to wrong. This work reframes the goal as 'selective trust': the model should judge each piece of context on its merits. To study it, the authors built a test set that shows the same question four ways — clean with no context, with a misleading hint, with a correct helpful hint, and with an irrelevant one — and a score that counts how often a bad hint sabotages an answer the model got right on its own. They then train the model on paired examples of these good-hint/bad-hint situations so it learns to lean in when context helps and shrug it off when it doesn't.

Technical view

The paper introduces MIST, a human-annotated benchmark rendering each reasoning item under four matched context conditions (clean, misleading, correct-context, irrelevant), and SC2W, a paired metric measuring how often a misleading signal flips a clean-correct answer to incorrect. A broad benchmark study finds this misleading-context susceptibility is essentially universal across models. Their method, SCOPE, mines clean-correct/misleading-wrong failure pairs and optimizes a standard DPO objective over these matched preferences to induce selective rather than blanket trust. Practitioners can reuse the four-condition rendering protocol and SC2W to audit context robustness in RAG or tool-augmented systems, and apply the DPO recipe on mined failure pairs to harden a deployed model without collapsing into context-ignoring behavior.

arXiv · cs.CLBuildable★ flagship

The Bitter Lesson of Tool Calling

Let AI call tools by writing code, not filling out rigid forms.

Modern AI agents 'use tools' — searching, calling APIs, running calculations — usually by emitting a strict form-like JSON message for each call. But models that can write code could instead just write a short script that calls those tools directly, naturally chaining several together or running them at once. This paper tests, head to head, whether the code approach actually works better across many different AI models on a standard agent benchmark. The tools are exposed as typed Python function stubs the model calls from code, with everything executed in a single turn. The finding echoes the field's 'bitter lesson': the more flexible, general code approach usually wins as models get stronger.

Technical view

The authors empirically compare programmatic tool calling (PTC) against native JSON tool calling across 14 language models on BFCL v4 under real-world task conditions. In PTC, tools are exposed as typed Python stubs invoked through generated code, with execution and result handling collapsed into a single agent turn, enabling natural chaining and parallelization. PTC matches or exceeds JSON tool calling in 11 of 14 models, with the GPT-5.6 family showing a 10.6% improvement. Practitioners building agent frameworks can adopt the typed-stub-plus-single-turn-execution pattern to boost tool-use accuracy on code-capable models, though the 3 regressions suggest the gain is model-dependent and worth A/B testing per deployment.

arXiv · cs.AIConceptual★ flagship

Tracing the Heart: An Evidence-Linked Pipeline for Heart-Failure Feature Engineering

An AI pipeline that turns messy heart-failure records into evidence-backed research features.

Medical researchers spend nearly half their time wrangling raw electronic health records into clean, usable variables — a huge, tedious bottleneck, especially for a complex condition like heart failure that spans scattered notes, labs, and guidelines. This work builds a team of cooperating AI agents that automatically pulls the relevant data together and computes clinically meaningful features, while keeping a traceable link back to the guideline or evidence that justifies each one. Crucially, every generated feature is checked for correct structure, scored against a clinical rubric, and audited for where its data came from, so a clinician can see why the number is what it is. They tested it on 500 synthetic patient records drawn from nine different record tables. The payoff is faster, more trustworthy, and more maintainable data preparation for clinical AI and studies.

Technical view

nMAS (Nimblemind Multi-Agent System) is an evidence-linked, rubric-grounded multi-agent pipeline for automated heart-failure feature engineering from fragmented EHR data. Evaluated on 500 dummy patient records spanning nine EHR source tables, it generated 132 structured and 70 rubric-scored aggregated features, each verified for structural integrity, rubric compliance, and provenance, with a restricted audit step. The design directly targets the maintainability and evidence-traceability gaps of prior rule-based and single-LLM approaches by grounding outputs in guideline-based rubrics and preserving data lineage. Teams building clinical feature stores could adapt the rubric-plus-provenance verification scaffold to make LLM-generated variables auditable, though real-world validation beyond synthetic records remains the next step.

arXiv · cs.CYConceptual

Investigating Artificial Intelligence Digital Sovereignty in Mobile Shopping Apps: A Case Study of Nigeria

Nigerian shopping apps quietly run AI on your data — and barely tell you.

Researchers looked at popular Nigerian e-commerce apps to see how much artificial intelligence they use behind the scenes — things like recommendation engines or fraud detection — and whether the apps actually tell users about it. They combined 'forensic' digging through the apps' code with reading company policy documents to see what AI features exist versus what's disclosed. The finding: AI is everywhere in these apps, but companies rarely explain it clearly, leaving users unaware of how much control they've handed over. This matters because 'digital sovereignty' — a country's or person's ability to control their own data and technology — erodes quietly when people can't see what's happening under the hood.

Technical view

The study applies an interpretive, mixed-methods design: static/forensic analysis of Android APKs to detect embedded AI features (e.g., ML libraries, recommendation logic, fraud models) cross-referenced with contextual document analysis of terms of service and privacy disclosures. Results show a gap between AI feature prevalence and transparency reporting, framed against Nigeria's socio-economic context of rising platform dependence and uneven digital literacy. Practitioners could replicate the forensic pipeline (APK decompilation, dependency/library fingerprinting) as an auditing methodology for regulatory or consumer-protection purposes in other emerging markets.

arXiv · cs.LGConceptual

An Optimal Agnostic PAC Algorithm

A math proof nails the exact best way to learn from noisy, imperfect data.

In machine learning theory, 'PAC learning' asks: given a bunch of examples, how many do you need to reliably learn a good rule, even when no rule is perfect (the 'agnostic' case, where some noise or overlap is unavoidable)? This paper builds a learning algorithm and proves a tight mathematical guarantee on exactly how its error shrinks as you feed it more data, matching a best-possible bound that theorists proved decades ago. It essentially closes a long-standing gap between what's theoretically possible and what an actual algorithm can achieve. It matters because it settles a foundational question about the true 'cost' of learning under uncertainty, giving other researchers a solid baseline to build on.

Technical view

For a hypothesis class H of VC dimension d, the authors construct a learner achieving risk L(ĥ) ≤ L* + O(√(L*(d+log(1/δ))/n) + (d+log(1/δ))/n) with probability 1-δ, matching the classical minimax lower bounds of Devroye–Györfi–Lugosi up to universal constants — resolving the exact sample complexity of agnostic PAC learning at every fixed noise level L*, not just asymptotically. This closes the gap between the 'fast rate' regime (L*=0) and the worst-case O(√(d/n)) regime by interpolating correctly for all L*. Practitioners in learning theory can use this as the definitive reference bound when analyzing new agnostic learners or when arguing optimality of empirical risk minimization variants.

arXiv · cs.GTBuildable

AV-AIVAT: 74x Cheaper Agent Evaluation with Certified Anytime-Valid Stopping in Imperfect-Information Games

Testing which AI poker bot is better, for 74x less money and compute.

To find out which of two AI agents is actually stronger at a game involving hidden information (like poker), you normally have to make them play thousands of rounds, which costs real money in compute or expert time. The problem is you never know in advance how many games you'll need — stop too early and the result is unreliable, stop too late and you waste resources. This paper builds a statistical method that watches the evidence as it accumulates and stops the very moment the answer becomes clear, without breaking the rules of valid statistics, combined with a variance-reduction trick that already made each game 54x more informative. Together it makes evaluating AI agents dramatically cheaper and faster while keeping the confidence guarantees mathematically sound.

Technical view

AV-AIVAT combines AIVAT — a conditional mean-zero control-variate correction that reduces variance in imperfect-information games, giving a median 54x variance reduction across 71,439 paired HUNL poker hands from 15 LLM agent configurations — with anytime-valid confidence sequences (CSs) that permit continuous monitoring and optional stopping without inflating the nominal error rate, unlike naive fixed-sample confidence intervals. The combined method yields a claimed 74x reduction in required games/cost versus fixed-budget evaluation while preserving statistical validity at every stopping time. Practitioners evaluating LLM or RL agents in adversarial/imperfect-information settings (poker, negotiation, auctions) could adopt this stopping rule to cut evaluation budgets without sacrificing rigor.

arXiv · cs.AIRunnable

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping

AI video models can't reliably count a bouncing ball's bounces past a dozen.

You'd think a state-of-the-art AI that watches video could easily count how many times a ball bounces or a person blinks — but this paper shows that's surprisingly hard for these models to do reliably. The researchers built controlled test videos (bouncing balls, blinking eyes, objects changing state) where they know the exact true event count and timing, so they can check the AI's answer against ground truth event-by-event, not just whether the final number happened to be right. They found a 'staged failure': models do fine with a few slow events but break down as the count or speed increases — for example, one top model tops out reliably counting only around 12 events. This matters because simple event-counting is a basic building block for AI systems meant to understand real-world video, like security footage or sports analysis, and current models aren't there yet.

Technical view

The authors introduce trace-grounded parametric profiling: synthetic video benchmarks (bouncing-ball wall contacts, blink detection, categorical state transitions) with independently controllable event count N and frequency F, paired with executable ground-truth traces enabling timestamp-level (not just final-answer) evaluation of counting accuracy. Testing across 2,190 videos reveals a 'staged temporal failure' — accuracy degrades predictably as N and F increase, with Gemini 3.6 Flash reliable up to only ~12 events for persistent state transitions at an 80% threshold. The controlled decoupling of event count/rate/duration from visual complexity isolates counting/bookkeeping as a distinct failure mode from general video understanding, giving practitioners a diagnostic benchmark to probe temporal reasoning limits in any video-language model.

arXiv · cs.GTConceptual

Resourced Authority A Mechanism-Design Model for Participatory Governance of Deployed AI Agents

A voting-and-budget system lets humans keep an AI agent on a leash after launch.

Once you deploy an AI agent, how do you keep ongoing control over it — not just at launch, but continuously? This paper designs a formal governance game: verified human stakeholders periodically contribute 'votes' (using a special currency separate from the AI's actual compute budget) to either approve or reject the agent's continued operation. Their contributions get combined into a weighted support score, and if support crosses a threshold (with some built-in stickiness so it doesn't flip-flop constantly), the system automatically grants or revokes the AI's authorization — enforced by literally controlling how much computing power it's allowed to use. The core idea is that compute is a practical lever for real, ongoing human oversight of AI, rather than a one-time approval that's forgotten after deployment.

Technical view

The paper formalizes participatory AI governance as a repeated extensive-form mechanism: stakeholders sequentially stake a distinct 'governance currency' in provision/rejection markets each period; a funding aggregator computes breadth-weighted effective support, and a two-threshold gate with hysteresis converts that support into a binary authorization signal. This authorization is coupled (via a bounded coupling map) to the agent's compute budget, making enforcement self-executing rather than reliant on post-hoc auditing — essentially treating compute allocation as the enforcement mechanism for a commons/compliance overlay on deployers. It's a mechanism-design contribution (game-theoretic, with equilibrium/incentive analysis implied) rather than an implemented system; researchers in AI governance or mechanism design could build on the threshold/hysteresis formalism for other resource-gated oversight schemes.

arXiv · cs.LGBuildable

CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal Tasks

An AI system auto-writes coding challenges precisely tuned to be 'hard but fair.'

When training AI agents to work in a terminal (writing and running code, essentially), you need practice tasks that are actually solvable but still push the agent's abilities — too easy and it learns nothing, too hard and it never succeeds enough to learn. CalibForge is a system that automatically generates and then adjusts these coding tasks by testing them against a pool of different AI 'solvers,' checking where the solvers disagree or where a strong solver succeeds but a weak one fails, and using that signal to calibrate task difficulty. Using this, they built over 5,000 tuned tasks, and experiments show tasks calibrated this way teach agents better than tasks that were just checked for basic feasibility. It matters because good training data is often the bottleneck for making AI agents better at real coding/ops work, and this automates a key part of curating it.

Technical view

CalibForge is an autonomous terminal-task synthesis pipeline that goes beyond executable-feasibility validation by using solver-relative calibration to place tasks in a 'learnable zone': multi-solver calibration exploits disagreement across a heterogeneous solver pool, while contrastive solver calibration targets tasks fitting a strong-pass/weak-fail pattern. The pipeline produced 5,431 calibrated terminal tasks, and ablations show both calibration strategies outperform authoring-plus-validation-only or single-solver baselines as training supervision. This gives practitioners building terminal/coding agents a reproducible recipe — solver pool + adversarial calibration loop — for scaling task datasets that stay appropriately challenging as agent capability improves.

arXiv · cs.AIConceptual

Challenges in Evaluating Explanation Methods for Static and Evolving Data

How do you know if an AI's 'explanation' of itself is actually any good?

AI systems increasingly come with explanation tools that claim to show why they made a decision (like flagging bias in an image classifier), but this paper argues we don't have good ways to check whether those explanations are actually trustworthy or useful. Using a real system called DetoxAI as a case study, the authors run a 'human-grounded' evaluation — meaning they check whether real people find the explanations helpful and accurate, not just whether they look plausible mathematically. They also tackle a harder problem: what happens to explanations when the underlying data keeps changing over time ('concept drift'), like trends shifting or new patterns emerging, using a technique called counterfactuals (showing what would need to change for a different outcome). It matters because as AI gets deployed in evolving, high-stakes settings like space applications, we need to trust not just the AI's answers but the reasons it gives for them, and reasons that go stale are dangerous.

Technical view

The paper surveys evaluation gaps in Explainable AI (XAI), grounding the discussion in DetoxAI, an image-recognition system for bias detection and concept unlearning, and presents a human-grounded evaluation study of image-classification explanation methods as a case example. It extends the discussion to explanation methods under concept drift in data streams, reporting experience adapting counterfactual explanations (which identify minimal input changes that would flip a model's decision) to non-stationary settings, and frames the open challenge of jointly tracking the co-evolution of data, models, and their explanations over time. This is a position/experience paper (workshop proceedings) rather than a new algorithm — useful for practitioners designing XAI evaluation protocols or maintaining explanation systems in production pipelines subject to drift.

arXiv · cs.CLBuildable

RP-OPSD: Reasoning-Pivot-Guided On-Policy Self-Distillation for Multilingual Reasoning Transfer

Teaching AI to reason in low-resource languages by copying just the key decision points.

Large language models are often much better at step-by-step reasoning in English than in less-common languages, because most training data is in English. One fix is 'self-distillation,' where a stronger 'teacher' version of the model guides a 'student' version on its own generated reasoning attempts — but existing methods treat every word of the reasoning equally, wasting effort on unimportant filler. This paper's insight is that only certain moments in a reasoning chain — 'pivots,' where the model decides to go one direction or another — actually matter for getting the answer right, so they focus the teaching signal there, comparing what the teacher would do with versus without seeing an English version of the solution as a hint. This lets a model transfer its reasoning skill into new languages more efficiently, which matters for making AI genuinely useful to the majority of the world that doesn't primarily use English.

Technical view

RP-OPSD builds on on-policy self-distillation (OPSD), where a teacher provides dense token-level supervision on rollouts sampled from the student itself, but reweights that supervision toward 'reasoning pivots' — tokens/decisions that redirect the reasoning trajectory — identified via the distributional shift between two teacher views of the same rollout, one conditioned on an English reference solution and one without. This creates an operational, unsupervised signal for locating transfer-critical tokens without needing explicit annotation of reasoning structure. The method targets cross-lingual reasoning transfer specifically, and practitioners building multilingual reasoning LLMs could adopt the pivot-detection mechanism as a drop-in reweighting scheme atop existing OPSD/distillation training loops.

arXiv · cs.AIBuildable

TRAJDEBUG: Tracing Error Lifecycle to Identify Critical Failures in Long-Horizon Agent Trajectories

When an AI agent's long plan goes wrong, this tool finds exactly which step first broke it.

AI agents that carry out long multi-step tasks (like browsing the web or writing code across many turns) often fail, but pinpointing which single step doomed the whole run is hard because the clue might be buried far back in a wall of past instructions and observations. TrajDebug tackles this by compressing the agent's history into digestible chunks at different levels of detail, then hunting through them for concrete evidence tying each step to the eventual failure. It also recognizes that a failed run can contain several small mistakes, only one of which actually caused the crash, so it has to distinguish the 'real' culprit from harmless side errors. The payoff is faster debugging of AI agents, similar to a flight recorder that tells engineers the exact moment that caused a crash instead of just 'something went wrong.'

Technical view

TrajDebug addresses critical-error identification in long-horizon LLM agent trajectories by combining multi-granularity history compression (summarizing distant context at varying resolution) with evidence-based localization that traces causal links between steps and the final failure. It explicitly separates locally-erroneous-but-harmless steps from the step whose downstream effects actually determine the trajectory's outcome, framing this as error-lifecycle tracing rather than single-step anomaly detection. Practitioners building agent evaluation or RL-from-failure pipelines could use this as an automatic critic generating step-level credit-assignment signals instead of relying on end-to-end reward alone.

arXiv · stat.MLBuildable

Scalable estimation of VARMA models

A math trick makes forecasting complex, interlinked time series (like markets) fast at any scale.

VARMA models are a classic statistical tool for forecasting multiple related time series at once (think interest rates, unemployment, and inflation moving together), and their 'moving-average' piece lets a model capture rich patterns with far fewer numbers than a plain autoregressive model needs. The catch has always been that fitting these models to real data was painfully slow and unstable as the number of series grew, since every tweak during fitting required re-scanning the entire dataset. This paper introduces a new way to reparametrize and fit VARMA models so each fitting step only touches a small, fixed-size summary of the data instead of the whole time series, using a mathematical shortcut (a Fourier-based identity) plus built-in guardrails keeping the model well-behaved. The result is that a previously 'impractical beyond a few variables' method becomes usable at real-world scale, which matters for forecasting large interacting systems in economics, climate, or engineering.

Technical view

The paper introduces a VARMA estimation framework whose per-iteration cost is independent of series length T, achieved via a partial-autocorrelation reparametrization guaranteeing stationarity/invertibility by construction, Gaussian priors with separate diagonal/off-diagonal scales for regularization, and loss functions expressed purely through fixed-size sufficient statistics computed via a Parseval (Fourier-domain) identity rather than direct likelihood evaluation over the full series. This removes the classical bottleneck of non-convex, non-identified VARMA likelihoods requiring O(T) work per gradient step. Practitioners can use this to fit high-dimensional VARMA models at scale where standard MLE/Kalman-filter approaches previously failed to converge or were computationally infeasible.

arXiv · stat.MLConceptual

Optimal Rates for Learning with Monotone Adversaries

Proves a sneaky data-poisoning trick genuinely makes learning harder — it's not just a bad algorithm.

Imagine a machine-learning student who gets a batch of correctly-labeled training examples, but an adversary sneaks in a few more genuine, correctly-labeled examples chosen to mislead the study process, then shuffles everything together so the student can't tell which is which. Earlier work showed the standard 'just minimize errors' learning strategy does worse in this setting than with pristine data, picking up an extra logarithmic penalty, and asked whether smarter algorithms could avoid that penalty entirely. This paper proves the answer is no: the penalty isn't a quirk of any particular algorithm but a mathematically unavoidable cost of the adversary's interference, for a wide class of problems, backed by a rigorous lower-bound proof. It matters because it tells researchers not to chase an impossible cleanup — the extra cost of learning from tampered-but-truthful data is fundamental.

Technical view

The paper resolves an open question from Larsen, Pabbaraju, and Shetty about learning under monotone adversarial insertion, where an adversary appends correctly-labeled points (dependent on the clean sample) before shuffling, breaking exchangeability. Prior work showed ERM achieves O((d/n)log(n/d)) error for VC dimension d, above the PAC-optimal Θ(d/n) rate, and asked whether the log factor is avoidable. This paper proves a matching lower bound showing the extra log(n/d) factor is information-theoretically inherent beyond a certain VC dimension threshold, not an artifact of specific algorithms, establishing tight optimal rates for this adversarial learning model.

arXiv · cs.DBBuildable

Tytan: Interactive Neurosymbolic Construction of Analytic Semantic Schemas from Relational Data

An AI that looks at your messy database and figures out, and asks about, what it actually means.

Before you can ask a chatbot questions about your company's data or auto-generate reports from it, someone has to explain what the tables and columns actually mean — which entities they represent, which numbers are measures versus which are just ID labels, and how tables relate to each other. Today that 'semantic layer' is written by hand by data experts, which is slow, expensive, and a bottleneck for anyone else wanting to use the data. Tytan automates this by combining rule-based analysis of the database structure with an AI language model's ability to guess sensible names and meanings, and — crucially — when the evidence is genuinely ambiguous, it asks the user a plain-English question instead of silently guessing wrong. This could let ordinary business users query and report on their own data without waiting on a data engineer to document everything first.

Technical view

TYTAN automatically constructs an analytic semantic schema (entities, measure/identifier column roles, table-join relationships defining analysis units) from a relational database, optionally combined with a short natural-language user description. It fuses symbolic analysis of schema/constraint structure with LLM-based inference for entity proposal, role assignment, and naming, falling back to interactive clarification questions when signals are ambiguous rather than committing to unreliable guesses. This targets the knowledge-acquisition bottleneck in BI/NL-to-SQL pipelines; practitioners could adopt the human-in-the-loop disambiguation pattern as a general technique for grounding LLM schema inference against ambiguous structured data.

arXiv · cs.CLBuildable

Benchmarking the Benchmarks: Evaluating Benchmarks for Conversational Agents

Who checks the test-makers? This tool grades AI chatbot benchmarks for being fair and thorough.

When companies build AI customer-service or task-completing chatbots, they measure quality using 'benchmarks' — standardized test scenarios — but nobody usually checks whether those benchmarks themselves are any good, and a sloppy benchmark full of inconsistent or overly simple tasks can give a false sense of how well an AI actually performs. This paper builds an automated grading system that uses AI judges to score a benchmark on three things: whether its tasks are internally consistent, how genuinely complex the scenarios are, and how much of the real range of policies and situations it actually covers. The team checked that their AI judges agree with human experts and tested the system on benchmarks of varying quality, including ones deliberately degraded, to prove it can tell good from bad. This is a check on the checkers — helping the industry trust that 'our chatbot scored 95%' actually means something.

Technical view

The paper proposes a reference-free evaluation framework that uses LLM judges to score task-oriented conversational-agent benchmarks along three axes: task consistency, scenario complexity, and policy coverage, producing actionable diagnostics rather than a single opaque quality number. Validation includes agreement with independent human annotations, differentiation across benchmarks generated by LLMs of varying capability, and sensitivity to controlled quality-degrading perturbations, plus applicability to manually-curated benchmarks. Practitioners building or selecting conversational-agent benchmarks can use this as an automated meta-evaluation tool to audit benchmark quality before trusting leaderboard results derived from it.

arXiv · cs.CLRunnable

Benchmarking and Enhancing LLMs for Rule-Intensive Review of National Standard Documents

Testing whether AI can proofread dense national technical standards like a compliance expert.

Countries like China publish huge, tightly-regulated 'national standard' documents (think building codes or manufacturing specs) that must follow strict rules about terminology, wording, and internal consistency — and checking a draft standard against all those rules today is slow, expert-driven work. This paper builds the first benchmark specifically for testing whether AI language models can do that kind of rule-heavy document review, breaking the task into a structured checklist covering document structure, whether content matches its stated scope, proper legal-style wording, and consistent terminology throughout. It's essentially a report card for AI models on a very specific, high-stakes reading task: catching subtle rule violations in long formal documents. This matters for any organization wanting AI to speed up compliance and quality-control review of regulations, contracts, or technical specs.

Technical view

GB/T-Bench is introduced as the first benchmark for structured rule-intensive review of national standard documents (China's GB/T standards), built around a hierarchical GB/T Review Taxonomy covering document structure, scope alignment, normative modality, terminology consistency, and cross-section normative rules. It targets a gap in existing LLM benchmarks, which focus on domain QA rather than intrinsic document-quality review requiring rule-grounded, long-context, structure-aware reasoning. The benchmark and taxonomy give practitioners a concrete testbed for evaluating and likely fine-tuning LLMs on regulatory/compliance-style document auditing, with direct applicability to other rule-governed document domains like legal or engineering standards.

arXiv · cs.CVRunnable

Does FLAIR super-resolution erase or hallucinate small white-matter lesions?

Does sharpening blurry brain scans with AI hide real damage or invent fake damage? They checked.

Doctors look for small bright spots on brain MRI scans (called white matter hyperintensities) that signal blood-vessel damage or early dementia, but scans taken in typical clinics are often blurry in one direction, making tiny lesions hard to see. AI-based 'super-resolution' tools promise to sharpen these blurry scans into clean 3D images, but nobody had carefully checked whether that sharpening might accidentally erase real small lesions or, just as dangerous, invent lesions that were never there. The researchers took 29 people's sharp, high-quality brain scans that had been carefully hand-labeled by a doctor, deliberately blurred them to mimic typical clinical scans, then ran several AI sharpening methods and compared the results against the hand-labeled truth. This matters directly for patient safety: a hospital adopting these tools without checking this risks missing real disease or triggering false alarms.

Technical view

The study evaluates whether super-resolution (SR) methods, applied prior to white-matter-hyperintensity (WMH) segmentation, preserve or distort lesion content, using 29 ADNI subjects with 1mm isotropic high-resolution FLAIR scans and expert manual WMH segmentations as ground truth, degraded to simulated 3mm and 5mm through-plane acquisitions. Three SR approaches were compared: a multi-contrast implicit neural representation (INR), a single-contrast self-supervised model (ECLARE), and simple cubic interpolation, assessed for their effect on downstream lesion detectability (erasure of true small lesions vs. hallucination of false ones). This gives clinicians and imaging researchers empirical evidence on SR method choice before deploying SR-then-segment pipelines in WMH burden quantification for cerebrovascular/neurodegenerative studies.

arXiv · cs.LGBuildable

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction

Fixes a mismatch so AI reward-judges can actually train other AIs via reinforcement learning.

When training an AI model with reinforcement learning (rewarding good outputs, penalizing bad ones), you need a 'reward model' to judge quality — and newer, more capable reward models work by comparing two answers and saying which is better, rather than giving a single numeric score. The problem is that standard reinforcement-learning algorithms expect a single score per answer, so these comparison-based judges haven't been able to plug in effectively, wasting their strength. This paper's fix, RRC, converts the AI judge's relative preferences ('this one's better than that one') into usable numeric rewards, either by comparing multiple AI-generated answers against each other or against a fixed reference answer. The result is training signals that actually match how these smarter judges naturally think, which should make AI systems trained this way improve faster and more reliably.

Technical view

RRC (Ranking-based Reward Construction) resolves a mismatch between generative reward models — which excel at comparative response ranking — and standard RL algorithms requiring scalar per-response rewards, a mismatch the authors identify as the reason generative RMs have underperformed their potential in RL fine-tuning. It offers two reward-construction strategies: self-competitive ranking (deriving rewards from pairwise/groupwise comparisons among sampled rollouts) and anchor-guided ranking (comparing against a fixed reference for scalable reward assignment). Practitioners doing RLHF/RLAIF with generative or LLM-judge reward models can adopt RRC as a drop-in reward-shaping layer within standard PPO/GRPO-style RL pipelines.

arXiv · cs.CVBuildable

UQ-Loc: Uncertainty-Aware LiDAR Scene Coordinate Regression

Robots can now say exactly how sure they are about where they think they are.

When a self-driving car or robot uses a laser scanner (LiDAR) to figure out its precise position, current systems just spit out a single 'best guess' location without saying how confident they are. UQ-Loc teaches the system to also estimate its own uncertainty for every tiny chunk of the 3D point cloud it sees, like a weather forecast giving you '70°F, plus or minus a margin' instead of just '70°F.' It does this by having the model predict a statistical spread around each guess, trained so overconfident or underconfident predictions get penalized. This matters because a self-driving car that knows it's unsure near a busy intersection can slow down instead of confidently driving into a wall.

Technical view

UQ-Loc extends the LightLoc scene-coordinate-regression architecture with a head that predicts a full anisotropic 3x3 positive-definite covariance matrix per voxel, trained via NLL loss plus a kNN spatial-smoothness regularizer. At inference, a modified SC2-PCR pose solver uses uncertainty-weighted seed scoring and a Mahalanobis-distance inlier test instead of fixed thresholds, and calibration is evaluated with Expected Calibration Error (ECE) rather than accuracy alone. This gives a route to uncertainty-aware localization a practitioner could adopt by swapping the regression head and solver in an existing LightLoc-based pipeline.

arXiv · cs.AIBuildable

Beyond Top-K: Replacing Black-Box Retrieval with Interpretable Agentic Operations

An AI that reads financial tables like an auditor, not a search engine.

Most AI document-search tools chop long reports into text chunks and match them to your question by similarity — fine for prose, terrible for financial statements where a number's meaning depends on a unit label (like 'lakh' vs 'crore') written many lines above it. The researchers show this breaks badly: on a 780-page government financial report, most content is table rows, and naive chunking routinely strips a number of the context that tells you if it's off by 100x. Instead of chunk-and-match, their system READ treats retrieval as a set of transparent, traceable operations, like a spreadsheet formula you can inspect, so you can see exactly which row, column, and header a fact came from. This matters for anyone using AI on financial statements or audits, where a wrong decimal place isn't annoying — it's a compliance disaster.

Technical view

The paper diagnoses standard chunk-embed-topK RAG as structurally unsound for table-dense financial documents, quantifying failure modes: 86.8% of lines are table rows, median unit-header distance is 13 lines, and 27-30% of numeric chunks lack fiscal-year headers even with a table-aware chunker. READ replaces black-box vector retrieval with interpretable agentic operations that explicitly resolve header/unit attribution rather than relying on embedding proximity. Practitioners building RAG over structured filings (10-Ks, audit reports, regulatory returns) can use the paper's diagnostic metrics as a test suite for their own chunkers before adopting an agentic-operation approach.

arXiv · cs.AIConceptual

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization

Testing whether an AI can rewrite its own toolkit to make itself smarter.

When you use an AI assistant, its intelligence isn't just the raw model — it's also the 'harness' around it: the instructions, tools, memory, and logic that guide it. This paper asks whether an AI can improve that harness itself, iterating on prompts and code the way a human engineer would, given feedback on how well it's working. They built HarnessOpt-Bench, a standardized test where an 'optimizer' AI is handed a starter harness, a budget of trial evaluations, and scored feedback, then must edit the harness and submit a better version. This matters because as AI agents get deployed everywhere, AI that can self-improve its own scaffolding, rather than needing humans to hand-tune every agent, could be a major lever for capability gains.

Technical view

HarnessOpt-Bench formalizes automated harness optimization as evaluation-guided search under expensive, stochastic evaluation: an LLM+coding-harness 'optimizer' receives a target agent's seed harness, graded feedback, and a fixed target-evaluation budget, then iteratively edits code/prompts/control-flow and nominates a final candidate that is scored independently. This provides a common protocol, rather than ad hoc harness-tuning reports, for comparing frontier LLMs on meta-level agentic engineering. A practitioner could use the benchmark's task/budget structure to evaluate their own optimizer loops or as a target to beat when building automated prompt/tool-orchestration search systems.

arXiv · cs.AIBuildable

Bias Analysis of L2 Speaking Assessment Systems Using Concept Activation Vectors

Checking if an AI grading your spoken English secretly judges your accent, not your skill.

AI systems increasingly grade spoken English tests for non-native speakers, and it's crucial their scores reflect speaking ability, not irrelevant traits like a test-taker's native language or age. But these graders run on complex 'black-box' neural networks where it's hard to see what they're actually keying on. This paper uses Concept Activation Vectors (CAVs), a technique that finds a direction inside the AI's internal 'thought space' corresponding to a specific human concept, like 'speaker's first language,' and checks how strongly the grader's score leans on that direction. They apply this to two real grading systems, a text-only one and one that also listens to audio, to see whether unwanted biases have crept in. This matters directly for fairness: a biased test could unfairly penalize test-takers based on where they're from rather than their actual skill.

Technical view

The work extends Concept Activation Vector (CAV) analysis, previously applied to feature-based L2 speaking graders, to two transformer-based systems: a text-only BERT grader and a multimodal Whisper-based speech-and-text grader. CAVs are computed as linear directions in activation space corresponding to interpretable concepts (e.g., L1, age) and used to measure how strongly a grader's score-relevant representations align with those directions, i.e., detecting concept leakage into the proficiency signal. This gives auditors of deployed speaking-assessment systems a concrete, activation-based fairness diagnostic that generalizes beyond feature-based models to end-to-end neural graders.

arXiv · cs.LGBuildable

On-Policy Self-Distillation without Any Supervision

A language model that teaches itself to fix its own mistakes, with zero outside help.

When training an AI model to get better after its initial training, most methods still need some outside signal: a correct-answer key, feedback from an environment, or guidance from a bigger, smarter model. This paper shows a model can improve just by listening to itself: it generates many attempts at a problem, and where most of its own attempts agree, it trusts that as a likely-correct 'pseudo-answer.' Then it uses its own best, most confident correct reasoning as a teacher to fix the exact spots where its longer, wrong reasoning went astray, like a student catching their own error by comparing a quick correct approach to one that veered off track. This matters because it points toward AI that keeps improving without more human-labeled data, bigger teacher models, or external verifiers, which are often expensive or unavailable.

Technical view

U-OPSD samples multiple rollouts per prompt, forms a pseudo-solution via majority vote under a self-consistency threshold, then conditions a teacher distribution on the shortest pseudo-solution and distills it into prefixes of the model's own longest incorrect completion, targeting correction precisely where the model is confidently wrong. Unlike prior OPD/OPSD methods, it requires no ground-truth labels, environment reward, or larger teacher model, relying entirely on self-consistency as the supervisory signal. Practitioners could replicate this recipe for post-training LLMs on reasoning benchmarks using only unlabeled prompts and the base model's own sampling, provided the domain admits a well-defined self-consistency criterion.

arXiv · cs.AIConceptual

QuanTiMedAI: Quantum-Enhanced Time-Series Model guided by Agentic AI for Cardiac Arrest Mortality Prediction

An AI research assistant plus quantum computing team up to predict cardiac-arrest survival.

For ICU patients after cardiac arrest, doctors want to predict who is likely to survive, but most prediction models only look at a snapshot from admission, ignoring how vital signs change minute-by-minute over the stay. QuanTiMedAI combines two cutting-edge ideas: an 'agentic' AI, a language model that actively reasons about which clinical features matter, like a research assistant hunting for useful patterns, picks out the most meaningful signals, and a small quantum-inspired model tracks how those signals evolve over time to predict mortality risk. This matters because more accurate, time-aware predictions could help ICU teams intervene earlier for deteriorating patients rather than relying on a first-day-only assessment.

Technical view

The system pairs an agentic LLM for clinically-informed feature discovery and selection with a compact quantum recurrent network that models temporal physiological trajectories over an ICU stay, applied to cardiac-arrest mortality prediction from EHR data. The claimed contribution is that LLM-guided feature selection improves downstream performance versus the static early-admission summaries used in prior mortality models. Reproducing this would require both an agentic feature-selection pipeline and a quantum or quantum-simulated recurrent architecture, making it a heavier lift than a standard classical time-series baseline to validate the claimed gains.

arXiv · cs.CLBuildable

NeSy-RAG: Neuro-Symbolic RAG for Explainable Question Answering

An AI that answers questions by writing logical proofs it can show its work for.

Normal AI question-answering systems that pull in outside documents (RAG) are often a black box: you get an answer but can't easily trace which sentence it came from or whether the in-between reasoning makes sense. NeSy-RAG fixes this by converting each retrieved chunk of text into small logic statements, using Prolog, a programming language built for formal reasoning, that capture true-or-false claims. It then combines the right statements into logical queries to answer your question, so every step is a traceable, checkable piece of logic rather than a fuzzy neural guess, and it can even notice when it's missing a fact about you specifically that it needs. This matters for trust: in high-stakes uses like legal or medical Q&A, being able to show exactly why an answer was given, and flag what's missing, beats a confident-sounding guess.

Technical view

NeSy-RAG synthesizes attributable Prolog modules from retrieved chunks by generating semantically meaningful predicates encoding Boolean claims, some dependent on user-provided facts, using joint natural-language/code embeddings to retrieve and compose predicates into Prolog queries for answering. A symbolic knowledge-gap detection mechanism identifies when required user-specific facts are missing from context, enabling targeted clarification rather than silent hallucination. This neuro-symbolic architecture is a concrete blueprint for attributable, verifiable RAG that a practitioner could adapt to any domain with clean claim structure, trading some flexibility for auditability versus pure vector-retrieval RAG.

arXiv · cs.LGBuildable

BaKron: Efficient Quantization with Kronecker-Factored Hessians

A faster math trick to shrink giant AI models without losing their sharpest details.

To make huge neural networks smaller and cheaper to run, engineers 'quantize' them, rounding their numbers to lower precision, but doing this carelessly hurts accuracy. Smarter quantization methods use the shape of the model's error landscape (the Hessian) to decide how to round each number so overall damage is minimized, and the best methods use richer, two-sided information about how an error in one part of the network affects others. The catch is that this richer information has been too slow to compute in practice. BaKron introduces a smarter algorithm for this exact calculation, using a divide-and-conquer strategy that keeps the same speed class as older, cruder methods while capturing the richer information. This matters because it could let developers shrink large AI models more accurately without paying extra compute cost.

Technical view

BaKron accelerates two-sided Kronecker-factored-Hessian-based adaptive rounding, as used in BoA and YAQA and extending GPTQ-style one-sided quantization, by combining anti-diagonal parallelism with recursive divide-and-conquer. This reduces total work from O(m²n²) to O(mn(m+n)) for an m×n weight matrix while keeping O(m+n) sequential steps, matching GPTQ's cubic scaling but with richer curvature information incorporated. This directly removes the computational bottleneck that previously made two-sided Hessian-aware rounding impractical at scale, letting practitioners substitute BaKron's solver wherever GPTQ-style rounding is used to get YAQA/BoA-quality quantization at comparable runtime cost.

arXiv · stat.COBuildable

Learning Latent Memory States from Longitudinal Athlete Monitoring Data

A new kind of compact, reusable 'memory snapshot' for tracking athletes over time.

This paper proposes a new way to summarize an athlete's recent training and health history into a compact table of numbers, instead of leaving raw sensor data scattered across time. The problem is that repeated measurements on the same person (like fatigue, sleep, or workload) are hard to store and reuse in a standardized way across different analyses. Their solution is a 'memory operator' that reads a masked window of recent history and compresses it into a small set of numbers plus uncertainty estimates, much like how other statistics (say, principal components) reduce complex data to a few reusable scores. The point isn't the AI model doing the compressing, but treating the resulting table itself as a first-class, reusable scientific object that can be stored, queried, and reanalyzed — validated here on a soccer-monitoring dataset.

Technical view

Introduces the 'Latent Memory Table' as a reusable statistical object analogous to PCA score matrices or random-effects tables, distinct in emphasis from the encoder that produces it. A memory operator (implemented as a Transformer) maps masked windowed longitudinal history to a finite-dimensional latent state with uncertainty, and collecting these states across subjects/time yields the table. Validation spans six properties — recoverability, personalization, temporal coherence, interpretability, stability, reusability — aggregated into a composite quality index Q, demonstrated on the SoccerMon athlete-monitoring case study. Practitioners could adopt this as a standardized embedding layer for downstream tasks like injury risk or workload modeling, reusable across a statistical workflow rather than model-specific.

arXiv · cs.LGBuildable

Surv-IPTB: An Attention-Based Model for Estimating Individual Probability of Treatment Benefit with Survival Data

An AI estimates each patient's personal odds that a treatment will actually extend their life.

Clinical trials usually report average treatment effects, but doctors really want to know if this particular patient will benefit. This work builds a model, Surv-IPTB, that estimates each patient's individual probability of living longer because of treatment versus not, using survival data where some patients' true outcomes are unknown because the study ended before an event occurred (called censoring). It works by comparing pairs of patients — one treated, one not — and turning that comparison into a classification problem, with an attention mechanism (a way of letting the model focus on the most relevant comparisons) weighing the evidence. For the uncertain, censored cases, it represents the probability as a range rather than a single number, being honest about what it doesn't know. This moves precision medicine closer to actually personalized treatment recommendations rather than population averages.

Technical view

Surv-IPTB reformulates individual probability of treatment benefit (IPTB) estimation as binary classification over pairwise comparisons between treatment and control cohort members, rather than modeling survival curves directly. Right-censored observations are handled via imprecise probability representations (interval-valued probabilities) instead of standard imputation or IPCW weighting. An attention mechanism with learnable query-key transforms aggregates pairwise evidence and jointly learns soft class probabilities for censored pairs. This gives a differentiable, trainable route to individualized treatment effect estimation on time-to-event data, extendable to other pairwise-comparison-based causal inference setups.

arXiv · cs.LGConceptual

The Tamed Subgradient Unadjusted Langevin Algorithm beyond Convexity

A smarter random sampler that stays stable while exploring jagged, steep, non-convex math landscapes.

In statistics and machine learning, you often need to draw random samples from a complicated probability distribution — imagine trying to randomly explore a bumpy, spiky terrain that isn't a simple smooth bowl. Standard sampling algorithms based on 'Langevin dynamics' (nudging a particle downhill plus adding randomness) can break down when the terrain has sharp kinks, grows extremely steeply, or has multiple valleys. This paper introduces SG-TULA, a method that works directly with the rough, kinked slopes (subgradients) and 'tames' overly large steps to keep the simulation stable, avoiding expensive smoothing tricks used elsewhere. The payoff is rigorous mathematical guarantees on how fast and accurately this sampler converges, even applied to messy real objectives like regularized model pretraining.

Technical view

SG-TULA is an explicit discretization of Langevin diffusion that operates directly on subgradients (for non-smooth potentials) with taming to control instability from superlinear gradient growth, avoiding smoothing-based approximations. The paper derives non-asymptotic Wasserstein-2 convergence bounds with explicit dimension- and inverse-temperature-dependent constants, improving on existing rates for subgradient Langevin methods, plus excess risk bounds for the associated non-convex optimization problem. Assumptions are verified with explicit constants for a regularized pretraining objective, suggesting direct applicability to sampling/optimization in deep learning regularization settings. Useful for practitioners needing provable sampling guarantees under non-smooth, non-convex, heavy-gradient-growth potentials rather than heuristic MCMC.

arXiv · stat.MLConceptual

Stochastic Dynamics on Persistence Diagram Space via Reinforcement Learning

Teaching an AI to gradually reshape the 'topological fingerprints' of data toward a chosen scientific goal.

Persistence diagrams are a mathematical way of summarizing the shape of data — like counting loops, clusters, or voids at different scales — used across science to describe structure. Most prior work treats these diagrams as frozen snapshots, but this paper wants to model how such shape-summaries could evolve randomly over time, like a stochastic process. Their approach uses reinforcement learning (an AI trained by trial-and-reward) to make small, topology-aware edits to a diagram step by step, forming a controlled random walk through 'diagram space.' They prove this random walk settles into a stable, well-behaved long-run pattern, and can be steered toward diagrams that match a scientifically meaningful target. This opens a path to simulating and generating evolving structural data, not just analyzing static snapshots.

Technical view

Proposes an RL framework in which persistence diagrams (PDs) evolve via topology-aware local edit operations, forming controlled Markov processes over PD spaces of variable cardinality. The authors establish irreducibility, aperiodicity, and geometric ergodicity conditions guaranteeing existence of unique stationary distributions over PD space — a nontrivial result given PDs' variable-dimensional, non-Euclidean structure. Objectives incorporating distributional targets steer the RL policy toward scientifically relevant topological configurations, effectively enabling controlled generative modeling in persistent homology space. This gives topological data analysis (TDA) practitioners a principled way to build simulators or generative models for evolving topological summaries, e.g., in dynamic physical or biological systems.

arXiv · cs.CVRunnable

TLNM: Externally Validated Tooth Detection, Numbering and Segmentation from Smartphone Photographs Using Mask R-CNN

A phone-selfie AI maps and numbers your teeth without needing a dentist's X-ray machine.

Good dental screening usually requires clinical X-rays or intraoral cameras that most people don't have access to, which limits preventive oral care worldwide. This work builds an AI model that instead works on ordinary smartphone photos of teeth, detecting, numbering, and outlining each tooth. It's built on Mask R-CNN, a well-known image-segmentation neural network, trained on about 1,272 labeled smartphone photos, with extra tricks added to correct for weird lighting/color casts from phone cameras and to reject detections that don't anatomically make sense (like a tooth in an impossible position). The system was tested not just on its own held-out data but also on a separate, independent external dataset, which is a stronger test of real-world reliability. The goal is enabling affordable, do-it-yourself oral health self-screening from a phone.

Technical view

Presents a customized Mask R-CNN pipeline for tooth instance detection, FDI/anatomical numbering, and segmentation from consumer smartphone photographs, trained on 1,272 annotated images. Two domain-specific mechanisms address smartphone-photo variability: a masked gray-world white-balancing algorithm correcting color casts, and an anatomically constrained detection layer enforcing structurally valid tooth arrangements to suppress false positives. Evaluation includes internal held-out testing plus independent external testing (a genuine generalization check rare in this space), alongside ablations. Practitioners building consumer oral-health screening apps could replicate this pipeline directly, given its external validation and off-the-shelf Mask R-CNN backbone.

arXiv · cs.AIConceptual

The Illusion of Visual Tool-Use: A Causal Audit of Thinking with Images

Checking whether AI's 'let me zoom into the image' habit is real reasoning or just theater.

Modern multimodal AI models can actively crop and zoom into parts of an image while answering questions, mimicking how a person might lean in to look closer. But researchers noticed these models often gain little from doing this, cost a lot more compute, sometimes zoom into irrelevant spots, and can even do worse than if they'd just answered directly. This paper asks: does the extra visual information the model 'sees' from zooming actually cause it to answer differently, or is the zooming mostly cosmetic? They build a formal cause-and-effect map of the process and run three levels of tests — comparing full tool-use vs. direct answers, deliberately feeding corrupted zoomed-in images, and swapping out a single zoom result mid-reasoning — to isolate what the visual evidence really contributes. The findings matter because they reveal whether flashy 'AI looks closer' behavior is genuine perception or an illusion dressed up as reasoning.

Technical view

Formulates multimodal LLM visual tool-use (crop-and-zoom operations) as a causal graph distinguishing observation-mediated effects from action-induced shortcuts, then audits it via interventions at three granularities: policy-level (tool-use vs. direct inference), trajectory-level (corrupting all rollout observations), and step-level (counterfactually replacing a single observation under a fixed prefix). Introduces 'Visual Evidence Gain,' a step-level causal estimand isolating the genuine contribution of returned visual crops to the final answer, separate from confounding by the reasoning trajectory itself. This gives a concrete methodology to diagnose whether agentic visual tool-use pipelines are doing real causal grounding or just adding token overhead — directly usable for auditing and redesigning thinking-with-images systems.

arXiv · cs.CVBuildable

OTLesMix: Wasserstein Barycenter and Optimal Transport Map for Synthetic Lesion Generation with Diverse Shapes and Locations

Uses transport-optimization math to conjure realistic new disease-lesion shapes for training medical AI.

Training AI to spot diseases in medical scans needs lots of labeled examples of lesions (abnormal tissue), but real datasets are limited and don't cover the full range of shapes and locations lesions can take. Existing tricks that mix or blend real lesion images together tend to produce a narrow, repetitive variety of synthetic lesions. This paper's method, OTLesMix, instead uses 'optimal transport' — math that finds the most efficient way to morph one shape or distribution into another — along with a concept called a Wasserstein barycenter (essentially an averaged 'in-between' shape) to blend real lesions into new, more diverse and realistic synthetic ones. The result is training data with a wider variety of lesion shapes and positions, which should help segmentation models generalize better to real patients.

Technical view

OTLesMix generates synthetic training lesions by computing Wasserstein barycenters and optimal transport maps between real lesion samples, interpolating shape and spatial location in a principled geometric way rather than simple pixel/label mixing. This targets the diversity bottleneck of prior mixing-based augmentation methods for medical image segmentation, which tend to reproduce limited lesion morphology and placement. The method plugs into existing segmentation training pipelines as an augmentation module, and practitioners can apply it directly to boost robustness on rare or underrepresented lesion shapes/locations.

arXiv · cs.LGConceptual

Hypothesis Testing with Conditional Queries: Learnability and the Value of Interaction

Working out exactly how much of an edge follow-up questions give a statistical test over fixed ones.

When you're testing a hypothesis — say, checking if two systems behave differently — you can either pick all your test questions in advance, or ask questions adaptively, using earlier answers to choose the next one. This paper studies exactly how much that adaptiveness helps, using a formal model of 'conditional queries' over a finite set of possible outcomes. They prove you can reliably tell two possibilities apart only when there's a real mathematical gap between them in how they answer conditional questions; without that gap, no amount of querying beats a coin flip. More surprisingly, they show you can convert a clever adaptive testing strategy into a non-adaptive one (all questions fixed upfront) that performs nearly as well, at a computable, moderate extra cost in the number of questions. This matters for designing efficient, trustworthy evaluation protocols, like AI benchmarks, where committing to tests in advance is often required for fairness or logistics.

Technical view

Analyzes conditional-query hypothesis testing over a finite outcome space of size N, proving learnability holds iff two distribution classes have positive separation in pairwise conditional probabilities, with worst-case error pinned exactly at 1/2 at any finite budget when that separation vanishes. The central result constructs, for any T-query adaptive policy and confidence parameter ρ, a randomized non-adaptive procedure using O(N²(T + log(1/ρ))) pair queries — chosen before any responses are observed — whose simulated transcript closely matches the adaptive policy's behavior. This gives a quantitative bound on the 'value of interaction' in query complexity, directly applicable to designing fixed-in-advance evaluation protocols (e.g., model benchmarking or property testing) that must approximate adaptive testing power.

arXiv · cs.LGBuildable

RxnCLF: Contrastive Transformation-Aware Reaction Foundation Model for Improved Reactivity Prediction

Teaching AI to sniff out which chemical reactions will actually work, from 1.7 million examples.

Chemists often want to predict how well a reaction will work — its 'yield' — before running it in a lab, but computers have struggled because there's so little labeled data and countless possible reaction combinations. This work builds RxnCLF, a system that learns from reactions by merging the 'before' (reactants) and 'after' (products) into one connected picture, called a condensed reaction graph, so the AI sees the actual transformation instead of two disconnected snapshots. It trains this understanding using a self-teaching technique (contrastive learning) on 1.7 million real reactions from a chemistry database called Pistachio, letting it learn patterns without needing humans to label every example. The payoff is a more reliable way to predict which reactions will succeed, potentially speeding up drug and materials discovery.

Technical view

RxnCLF is a self-supervised contrastive learning framework for reaction representation, built on a condensed reaction graph (CRG) that unifies reactant and product atoms/bonds into a single graph rather than encoding them as disconnected structures, explicitly capturing the reaction-center transformation. It is pretrained on 1.7 million Pistachio reactions to learn a compact, continuous latent embedding space transferable to downstream reactivity/yield prediction tasks. This addresses known limitations of string- (SMILES), fingerprint-, and graph-based reaction encodings that only partially capture transformation semantics. Practitioners could fine-tune the pretrained embeddings on smaller labeled yield datasets for low-data reaction optimization tasks.

arXiv · cs.CVConceptual

MASS: Multiplayer World Models with Authoritative Shared State

A game-engine trick lets AI-generated worlds stay consistent no matter which player is looking where.

When AI video-generation models try to create shared virtual worlds for multiple players, they usually mix up 'what the world actually looks like' with 'what this specific camera sees,' causing wasted computation and glitches where different players see inconsistent versions of the same scene. MASS fixes this by borrowing an idea from multiplayer video games: keep one single, official copy of the world's state, and generate each player's view from that shared source on demand. A learned 'Logic Engine' updates this master state automatically based on everyone's actions (no manual programming needed), while a separate 'Rendering Engine' paints the picture for whichever camera asks. This split keeps everyone's view of the world honest and consistent, which matters for believable multiplayer games, simulations, or virtual training environments.

Technical view

MAS(S) disentangles world-model architecture into two learned components: a Logic Engine that recurrently advances a global, typed, authoritative state from joint multi-agent actions (replacing hand-written transition functions and serving as the sole memory/synchronization point), and a Rendering Engine that generates independent, on-demand views per requested camera from that shared state. This separation avoids entangling view-dependent visual latents with world dynamics, the source of redundant compute and cross-view inconsistency in prior multi-view video world models. The paper reports superior state accuracy and lower cross-view inconsistency versus state-of-the-art multi-view baselines. This architecture is directly relevant to anyone building scalable multi-agent/multiplayer generative simulators or game-world models.

arXiv · cs.LGBuildable

MetaboLLM: a metabolomics-specialized large language model for biochemical knowledge integration and predictive metabolite graph construction

A chatbot that's read all of metabolism science and turns patient chemistry into disease predictions.

Metabolomics — the study of small molecules in the body like sugars and fats — has knowledge scattered across many separate databases, making it hard for computers to use it for predictions. MetaboLLM is an AI language model specially trained on this scattered biochemical knowledge, refined through extra training stages and given the ability to look things up as needed. A companion tool, MetaboLLM-GIN, takes the descriptions this AI generates and turns them into graphs — network-style maps of how metabolites relate — which a second AI then uses to predict things about individual patients. In tests, it beat both general and medically-tuned AI models at understanding metabolomics, and its patient predictions were notably accurate for things like blood sugar spikes after heart surgery and hormone therapy classification, suggesting real clinical usefulness.

Technical view

MetaboLLM is built via continual pretraining, supervised fine-tuning, and structured retrieval augmentation on metabolomics literature/knowledge bases, evaluated across four backbone LLM families against base and medically-adapted models on knowledge, relational, and description benchmarks, with transfer to an external public benchmark. MetaboLLM-GIN converts the LLM's generated biochemical descriptions into metabolite graphs consumed by a graph isomorphism network (GIN) for patient-level classification, achieving AUCs of 0.8616 for stress hyperglycemia after coronary artery bypass grafting and 0.8123 for postmenopausal hormone-regimen classification, outperforming conventional models. This offers a template for LLM-to-graph pipelines in other biomedical omics domains where domain knowledge is fragmented across text sources.

arXiv · cs.CVRunnable

Toward Deployable Bangla Sign Language Recognition with Expert-Validated Data and a Lightweight Attention-Based Model

A tiny 300K-parameter phone app could finally read Bangla Sign Language accurately in real classrooms.

Deaf and hard-of-hearing people in Bangladesh mainly communicate using Bangla Sign Language, but existing computer recognition systems were trained on artificial lab data and rely on huge AI models too heavy for phones. The researchers built RSBdSL38, a new dataset of nearly 11,000 real, expert-checked photos covering all 38 hand signs (representing the 51-letter Bangla alphabet), collected directly from signers at special-needs schools. They then designed a lightweight AI model — just 298,470 parameters, tiny compared to typical vision models — using efficient building blocks like attention (letting the model focus on the most informative parts of an image) and multi-scale hand-shape detection. Trained from scratch, it correctly recognized signs 96.37% of the time, showing that accurate, phone-friendly sign language recognition is achievable and could genuinely widen access to education and services.

Technical view

The paper contributes RSBdSL38, an expert-validated dataset of 10,874 real-world images spanning all 38 BdSL hand signs (covering the 51-letter Bangla alphabet) collected from actual signers at three special-needs schools, addressing prior reliance on controlled-setting, unverified data. It pairs this with a lightweight attention-based CNN (298,470 parameters) combining grouped bottleneck residual blocks, channel/spatial attention, a multi-scale depthwise hand-feature block, dual pooling, and Swish activations — designed for on-device deployment rather than fine-tuning heavyweight pretrained backbones. Trained from scratch, it achieves 96.37% accuracy (95.72% ± 0.54% across five seeds), positioning it as a practical baseline for mobile/edge BdSL recognition apps. Practitioners could directly reuse the dataset and architecture for on-device sign language recognition research.

arXiv · stat.MLConceptual

Minimax Optimal Early-Stopped Gradient Descent for Gaussian Mixture Classification

Stopping AI training early, at exactly the right moment, turns out to be provably the best strategy.

When a computer learns to classify data (like sorting points into two groups) using an overly flexible model, it can perfectly separate even messy, overlapping data — but the resulting boundary, found by running the training process (gradient descent) to convergence, is often not actually the smartest one statistically. This paper studies a cleaner mathematical setting — data drawn from two overlapping Gaussian 'blobs' with some mislabeled points — and proves that if you stop the training at just the right time instead of letting it run forever, you get the mathematically best-possible classifier. They show this holds even as the complexity of the data structure decays in different ways (fast or slow). It's a theoretical result, but it explains why 'early stopping,' a trick practitioners already use somewhat informally, has a rigorous justification for being genuinely optimal, not just a heuristic hack.

Technical view

In overparameterized logistic-loss classification, GD diverges in norm but converges in direction to the max-margin interpolator, whose implicit bias is known to be statistically suboptimal in some regimes. The authors prove that for a Gaussian mixture model with label-flipping noise, GD stopped at an oracle-determined time achieves minimax-optimal excess 0-1 risk across covariance spectra with fast/continuous decay (polynomial and exponential), via a sharp upper bound on the early-stopped iterate matched to a lower bound over arbitrary classifiers. A key technical tool is a new calibration argument bridging the early-stopping iterate's behavior to the minimax rate, validated experimentally. This gives theorists a rigorous target for early-stopping schedules and could inform principled stopping-time selection rules in practical overparameterized classifiers.

arXiv · cs.LGConceptual

A Six-Dimensional Taxonomy of Post-Training Adaptation Techniques with Applications in AI Governance

A map of every way to tweak a trained AI model, built so regulators can finally compare them.

After an AI model is initially trained, developers often modify it further — retraining it, fine-tuning it for a task, teaching it new facts, making it forget things, or adjusting how confident it sounds — but these techniques are described inconsistently across research papers, making it hard to know what was actually done to a given model. This survey organizes all of these post-training tweaks into a clear six-part framework, sorting them by how they work, what goal they serve, how much data they need, how permanent the change is, how much of the model they touch, and what kind of model they apply to. It carefully separates terms people often confuse, like 'fine-tuning' versus 'giving the model a lookup tool' (retrieval augmentation) versus just 'prompting' it differently. The point is practical: policymakers and auditors need a shared vocabulary to describe and regulate how AI models have been changed after their initial training.

Technical view

This survey proposes a six-dimensional taxonomy for post-training adaptation techniques — spanning retraining, fine-tuning, parameter-efficient adaptation (e.g., LoRA-style methods), alignment, retrieval augmentation, model editing, unlearning, calibration, and multimodal instruction tuning — organized along mechanism, goal, data requirement, persistence, structural scope, and model type. It disambiguates commonly conflated terms (fine-tuning vs. retrieval augmentation vs. prompting) and traces how adaptation strategies evolved from classical ML through deep learning to foundation models. The framework is explicitly positioned for AI governance use cases, giving auditors/regulators a standardized vocabulary to describe how a deployed model diverges from its base training. Useful as a reference taxonomy for anyone writing model documentation, compliance reports, or comparative method surveys.

arXiv · cs.AIBuildable

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models

Teaching a student AI by weighting its mistakes based on how their errors have been trending.

AI reasoning models are often improved using reinforcement learning with rewards that only say 'right' or 'wrong' at the very end of a long chain of reasoning — a very sparse signal. A better trick, called on-policy self-distillation, has a stronger 'teacher' AI grade the student's reasoning at every step along the way, giving denser feedback. This paper points out that current versions of this trick treat every disagreement between student and teacher the same way, ignoring whether that disagreement is part of a long streak of similar mistakes or a one-off blip. DASH adjusts how much weight to give each moment of disagreement based on this history, essentially paying more attention to persistent patterns of confusion rather than treating every hiccup equally, which should make the training signal more informative and the resulting reasoning model better.

Technical view

DASH extends on-policy self-distillation (OPSD) for RLVR-trained reasoning LLMs by making the token-level supervision coefficient divergence-adaptive rather than uniform: instead of weighting every local teacher-student KL divergence identically regardless of position, DASH conditions the weighting on the local discrepancy's history within the rollout, recognizing that identical divergence magnitudes can reflect different trajectories of teacher-student mismatch. This targets underexploited temporal structure in dense token-level distillation signals derived from querying a privileged teacher at student-visited prefixes. The approach is directly applicable to anyone running teacher-guided RLVR pipelines for reasoning models, as a drop-in reweighting scheme atop existing OPSD implementations.

arXiv · cs.CVBuildable

PRISM: Distribution-Gated Flow Matching for Controllable Unpaired Image Translation

An AI image translator with a smart per-pixel dial deciding exactly what to keep and what to change.

Translating an image from one style or domain to another (say, turning a photo into a painting) without paired before/after examples requires deciding, pixel by pixel, what to preserve (like the subject's identity) and what to alter (like the style). Older methods use one single global knob for the whole image, which can't tell 'content to keep' apart from 'appearance to change' region by region. PRISM instead learns a gate for each part of the image, based on how far that part's features are from what a typical 'target style' image looks like — parts far from the target are allowed to change freely, while parts already close to the target style are protected. This same gate controls both the starting point of the AI's generation process and the pacing of its internal transformation steps, giving much finer, more controllable image editing without needing matched training pairs.

Technical view

PRISM is a GAN-free flow-matching model for unpaired image-to-image translation that replaces global noise/guidance-strength control with a learned per-feature spatial gate, computed from each source feature's standardized distance to the target feature distribution — features distant from the target distribution are freed to change while target-consistent features are preserved. This single gate jointly governs (1) initialization, by mixing the real source latent with a task-matched corruption, and (2) transport timing during the ODE integration used for flow-matching generation. By deriving both controls from the same distribution-aware signal rather than a single global hyperparameter, PRISM aims for finer per-region content/appearance disentanglement than prior diffusion-based unpaired translators. Practitioners working with flow-matching image translation pipelines could adopt the per-feature gating mechanism as a replacement for global guidance scales.

arXiv · cs.CVBuildable

Depth-Guided Video Object Counting in Crowded Scenes

Teaching cameras to count crowds accurately by adding a sense of depth.

This project is about getting AI to count things in a busy video—like people in a crowd or animals in a herd—even when they're overlapping and blocking each other. Regular systems rely only on color images, which struggle when objects pile up in front of each other. The trick here is adding depth information (basically a 3D sense of how far away each object is, like how our two eyes judge distance) and combining it with the color image so the system can tell separate but overlapping objects apart. They also built a system to avoid double-counting the same object as it moves across video frames. This matters for things like crowd safety monitoring, wildlife counts, or retail analytics where miscounting overlapping objects is a real problem.

Technical view

The authors propose DG-Det, a depth-guided detector that fuses RGB and depth via multi-scale cross-attention and adds explicit occlusion prediction to improve instance discrimination in crowded, occluded scenes for text/visual-prompted object counting. A unified de-duplication post-processing pipeline removes cross-frame redundant counts, addressing the temporal double-counting problem common in video counting. They also release a new RGB-D video object counting dataset with multi-category annotations per sequence, enabling depth-aware benchmarking. Practitioners could build on this by adapting the RGB-D cross-attention module to other dense detection/counting tasks where occlusion is a bottleneck.

arXiv · cs.CVBuildable

EmoWorld: A Decoupled Affective Field for Controllable Emotional Video Generation

A video AI that lets you dial up 'sad' or 'tense' like adjusting a thermostat.

AI video generators can be told what to show, but controlling the emotional mood of a generated video—separate from its content—has been hard because 'happy' or 'eerie' gets tangled up with the literal scene description. EmoWorld splits emotion into three separate dials: the overall visual atmosphere (lighting, color mood), the emotional meaning baked into the scene's semantics, and how that emotion builds or shifts over time. It works by first extracting reusable 'emotion directions' from pairs of neutral versus emotionally-edited images, then nudging the video-generation process along those directions during creation, without retraining the whole model. This kind of fine-grained, independent control could make AI-generated film, ads, or games far more expressive and directable.

Technical view

EmoWorld decouples atmosphere, semantic affect, and temporal progression in a frozen flow-matching Video DiT by pre-extracting layer-specific affect directions and a reusable cue library from geometry-preserving neutral/emotion-edited panorama pairs. At inference it applies three steering mechanisms: Visual Atmosphere Steering (injecting directions into hidden states), Semantic Affective Steering (a separately scalable prompt residual), and Temporal Affective Steering (interpolating endpoint residual fields across denoising and video time). On Wan2.2 it reports a 19% improvement in target-emotion alignment and 48% reduction in a temporal-fluctuation proxy, suggesting a training-free, composable steering approach practitioners could adapt to other diffusion transformer video models.

arXiv · cs.AIBuildable

TS-RAG: Retrieval Augmented Generation for Time Series Forecasting

Forecasting the future by having the AI look up similar patterns from the past first.

Time series forecasting means predicting what comes next in a sequence of numbers, like stock prices or sensor readings. Large language models got a boost from 'retrieval-augmented generation,' where the AI looks up relevant snippets of text before answering—TS-RAG asks whether the same trick works for numeric forecasting: retrieve similar past sequences as reference examples before predicting. The catch is that time series models are usually much smaller and simpler than language models, so just pasting reference data into a prompt like you would with text doesn't work well. TS-RAG designs a way to properly incorporate those retrieved similar sequences so the model can actually use them to improve its predictions. This matters because better forecasting with less need for massive training data could help smaller, cheaper models perform competitively.

Technical view

TS-RAG adapts retrieval-augmented generation to time series forecasting by retrieving similar historical sequences as references, but avoids naive prompt-concatenation (effective for LLMs but not for smaller time series models) in favor of a more integrated fusion mechanism tailored to transformer-based forecasters with limited parameters and training data. The core claim is that principled retrieval integration—rather than simple context-stuffing—yields accuracy gains for models lacking LLM-scale generative capacity. This is directly buildable: practitioners with existing time-series transformers could add a retrieval index over historical windows and experiment with the paper's fusion approach as a lightweight accuracy boost.

ROB

Robotics

20 new
arXiv · cs.ROConceptual★ flagship

$ω$-0: A Latent Predictive World Action Model for Concurrent Humanoid Loco-Manipulation

A humanoid robot that walks, balances, and manipulates objects as one fluid motion.

Household tasks often require a robot to do several things at once — step toward a shelf, shift its posture, keep its balance, and grab an object — all as one smooth coordinated act. Most humanoid robots instead treat walking and grabbing as separate modules bolted together, which makes combined tasks clumsy. This system, ω-0, takes a spoken instruction plus what the robot currently sees and feels, and directly outputs whole-body motions the robot can execute. Instead of trying to predict future camera video (expensive and noisy), it learns a compact 'mental preview' of what it will observe next, and uses that foresight to guide a whole-body motion generator. This lets one model handle moving and manipulating together on a real humanoid.

Technical view

ω-0 is a latent predictive whole-body world-action model that maps a language instruction, current visual observation, and proprioceptive state directly to controller-compatible whole-body action latents for real-robot execution. Rather than reconstructing future frames, it learns compact future-observation embeddings as a lightweight predictive objective, coupling latent visual foresight with a diffusion-based whole-body action generator. This unifies locomotion and manipulation into a single coordinated policy, addressing the arm-centric or video-centered limitations of prior world-action models. Robotics practitioners can build on the latent-foresight-plus-diffusion-action design to avoid costly video prediction while retaining predictive grounding for concurrent loco-manipulation, using egocentric observation as input.

arXiv · cs.ROConceptual★ flagship

DyPES-VLA: Learning Shared Dynamics Priors and Embodiment-Specific Control for Cross-Embodiment Manipulation

One robot brain that transfers manipulation skills across differently-shaped robots.

Vision-Language-Action models let robots take a camera view plus an instruction and produce actions, but training a single policy that works across many differently-built robots (different arms, grippers, bodies) is hard. Two problems block it: the models don't fully exploit the physics of interaction that's shared across all robots and data, and they need lots of manual work to reformat each robot's actions into a common language. DyPES-VLA tackles both by first teaching the vision-language model to predict how scenes will change — how objects move, make contact, and get pushed — so it captures shared physical dynamics regardless of which robot is acting. Then a separate, body-specific component translates that shared understanding into the particular robot's controls. This splits 'what physically happens' from 'how this specific robot does it,' improving transfer.

Technical view

DyPES-VLA separates shared Dynamics Priors from Embodiment-Specific control for cross-embodiment manipulation. The shared prior is learned by training the VLM with a future-prediction objective over cross-embodiment data, driving a shared query representation to encode object motion, contact, and interaction-induced scene changes. An embodiment-specific control module then maps these shared representations to per-robot actions, removing the need for extensive manual action normalization into a common format. Practitioners can adopt the future-prediction pretraining to build embodiment-agnostic dynamics features, then attach lightweight per-embodiment heads, improving generalist-policy transfer across heterogeneous robots without heavy preprocessing.

arXiv · cs.ROConceptual★ flagship

A Master-Salve Robot Manipulator for Needle-Based Teleoperation in MRI Chamber

A fluid-powered robot that lets a doctor steer a needle inside an MRI scanner.

MRI gives beautiful real-time images but its powerful magnet forbids ordinary metal motors and electronics near the patient, making it hard to perform needle procedures while scanning. This device is a master-slave robot: the doctor moves a handheld controller (the 'master') and a matching robot beside the patient (the 'slave') mirrors the motion, with force transmitted through fluid-filled tubes instead of magnetic motors, so it's MRI-safe. It handles both angling the needle and pushing it in, transmitting forces below one newton and motions below a millimeter faithfully over the long piping needed to reach outside the magnet. A digital controller adds modes beyond simple mirroring, allowing manual, digital, hybrid, and collaborative control where doctor and computer share the task. The result is precise, MRI-guided abdominal needle interventions in real time.

Technical view

The system is an MR-safe 2+1-DoF master-slave manipulator for abdominal interventions inside the MRI bore, transmitting motion and force via fluid transmission. High-input-impedance, low-leakage elastomeric fluid actuators handle remote angulation, while low-friction graphite piston cylinders drive the needle-insertion axis with sub-newton force transparency and sub-millimeter motion transmission over bedside piping lengths. A digital master controller adds multimodal control (manual, digital, hybrid, collaborative) beyond typical split-axis or mode-switchable hybrid configurations. Researchers in MR-guided robotics can build on the actuator delegation strategy — elastomeric actuators for angulation, graphite cylinders for insertion — to achieve force-transparent teleoperation compatible with real-time MRI guidance.

arXiv · cs.ROBuildable

GeniWorld: A Generalizable Interactive World Model for Robotic Manipulation via Visual Actions

A simulated 'dream world' where robots can practice actions before trying them for real.

Robots that manipulate objects need lots of trial-and-error to learn, but testing in the real world is slow and expensive, and simulations often don't generalize to new situations. GeniWorld builds an AI 'world model'—essentially a video-generating simulator that predicts what happens next when a robot takes an action—so robots can rehearse and be evaluated safely and cheaply. Its key trick is converting numeric robot commands into visual representations of the robot's actual arm shape (using standard robot-description files) so the model understands actions as spatially grounded, not just abstract numbers. It also separates 'how the robot's body moves' from 'how the environment behaves,' which helps it generalize to scenes it's never seen. This kind of world model could make training and testing general-purpose robots far cheaper and safer.

Technical view

GeniWorld builds an action-conditioned interactive world model atop pretrained video generative models, using URDF-based rendering to convert numerical robot actions into visual action representations for spatially grounded control, rather than conditioning on raw action vectors. By explicitly decoupling embodiment kinematics from environmental dynamics, the model reduces scene overfitting and improves generalization to out-of-distribution scenarios, addressing a known weakness of prior action-conditioned world models. This is relevant to robotics practitioners building simulation-free policy evaluation or model-based RL pipelines, since a generalizable visual world model could substitute for costly real-world rollouts.

arXiv · cs.RORunnable

Acoustic-driven millimetric helical robot: ultrasonic synergistic manipulation in confined fluidic environment

Tiny screw-shaped robots steered through your body using nothing but sound waves.

Imagine a robot smaller than a grain of rice, shaped like a tiny screw, that a doctor could steer through blood vessels or tight spaces inside the body without touching it—just using sound. This research shows how combining two different acoustic effects (the physical push of sound waves, and the swirling flow currents sound creates in fluid) lets these millimeter-sized helical robots move much more effectively than either effect alone. The team tested this both in computer simulations and real experiments, showing the robots could navigate flat surfaces, climb inclines, and even move straight up, all controlled by ultrasound. They also demonstrated semi-autonomous navigation, meaning the robot could follow a path with some independence rather than being manually steered every second. This kind of technology points toward future minimally invasive medical procedures like targeted drug delivery inside confined body spaces.

Technical view

The paper introduces a coordinated multi-acoustic-field strategy that combines acoustic radiation forces with acoustic streaming flows to drive millimeter-scale helical robots, overcoming the propulsion inefficiency that previously limited acoustic actuation to micro/nanoscale objects in confined fluidic (likely biological) environments. Multiphysics simulations modeled the coupled dynamics, and experiments validated planar navigation, inclined climbing, vertical motion, and semi-autonomous path-following enabled by the synergistic ultrasonic fields. This suggests a viable actuation mechanism for non-contact, non-invasive millimetric robotic navigation in confined channels (e.g., vasculature or catheterized spaces), with clear extension paths for closed-loop autonomous control and in vivo validation.

arXiv · cs.ROBuildable

In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use

Robots that listen to instructions instead of talking to themselves before acting.

Vision-Language-Action models are AI systems that let robots see a scene, read an instruction, and perform a physical task. A popular idea has been to make robots 'think out loud' in text (chain-of-thought reasoning) before acting, the way chatbots do—but this paper finds that actually hurts robot control: the reasoning isn't grounded in reality, it slows the robot down, and training the robot to both narrate and act pulls it in conflicting directions, so it ends up better at talking than doing. Instead, the authors argue robots should be good at understanding language given to them, not generating their own commentary. Their method trains the robot to consume grounded language input more effectively via extra training steps and by letting it call external tools when needed. This reframes how reasoning should be integrated into robot AI, favoring comprehension over self-narration.

Technical view

The paper empirically and analytically shows that free-form textual chain-of-thought degrades VLA control because the generated reasoning is ungrounded, its generation latency breaks closed-loop timing, and jointly optimizing reasoning and action tokens creates conflicting objectives that bias the policy toward narration over action. Their proposed framework instead endows the VLA with language *comprehension* competence via in-context post-training and agentic tool use, rather than free-form generation. This suggests a concrete design principle for VLA researchers: avoid training policies to emit CoT text jointly with actions, and instead invest reasoning capacity in grounded instruction-following and external tool invocation.

arXiv · cs.ROBuildable

Near-sensor Computing for Rapid Visuotactile Perception

A touch sensor chip that feels an object's shape in about a fifth of a millisecond.

Robots that use touch sensors (visuotactile sensors) to feel object shapes usually have to send that raw data to a separate computer for processing, which adds delay and eats power—like taking a photo and mailing it out to get developed before you can react. This research builds the processing directly into hardware right next to the sensor, so the shape-reconstruction math happens in a dedicated streaming circuit rather than a general-purpose computer. The result is extremely fast, predictable response times (a fraction of a millisecond) and low power use, because the circuit doesn't need to loop or branch depending on the data—it just streams through in one steady pass. Tested across 15 different contact shapes, the results closely matched more precise, slower reference calculations. This kind of near-sensor computing could let robots react to touch almost instantly, useful for delicate manipulation or fast reflexive grasping.

Technical view

The authors implement a fully streaming hardware pipeline for a spectral Poisson solver that reconstructs dense contact geometry from visuotactile surface-gradient measurements directly near the sensor, avoiding host-based processing latency and power overhead. The core logic runs at 166 MHz, consumes an estimated 347 mW, and produces the first depth value of a 128x128 frame in 35,107 cycles (0.211 ms fixed latency) with deterministic timing due to the absence of data-dependent branching or iterative convergence. Across 15 tested contact geometries, reconstructed depths closely matched double-precision reference computations, indicating the streaming architecture preserves accuracy while enabling real-time, low-power tactile sensing for robotics — a template for hardware practitioners looking to move dense-perception math off the host CPU/GPU and onto dedicated near-sensor silicon.

arXiv · cs.ROBuildable

ATP: Anatomical Torque with Passivity-based Control Framework for Safe Upper-Limb Exoskeleton Assistance

An exoskeleton that senses your muscles' natural forces to help your arm safely, for any movement.

Exoskeletons are wearable robotic devices that assist human movement, and most 'smart' assistance research has focused on legs, where movements are repetitive and predictable (like walking). Arms are much harder because they move in complex, non-repeating ways, making it tricky to know exactly how much force to apply without being unsafe or unhelpful. This paper trains a virtual muscle-control AI, using reinforcement learning in a detailed simulated model of human musculoskeletal anatomy, to figure out the natural torque (rotational force) the body would generate for any arm movement, without needing complex real-time biomechanical calculations. It then adds an online adjustment system that refines this estimate for the specific movement happening and keeps the physical interaction stable and safe (avoiding jerky tendon-related issues). This could lead to upper-limb exoskeletons that assist naturally and safely across a much wider range of everyday arm movements, useful for rehabilitation or physical labor support.

Technical view

ATP combines a reinforcement-learning-trained unified muscle controller, learned within a scalable musculoskeletal simulation framework, to generate anatomical reference torques for arbitrary upper-limb movements without explicit biomechanical modeling at runtime, with an online torque-refinement scheme that adapts these references per movement and suppresses tendon-induced instabilities via passivity-based control for safety guarantees. This addresses a gap versus lower-limb exoskeleton assistance, where periodic, weight-bearing gait patterns make torque estimation comparatively easier than for nonperiodic, high-DOF upper-limb tasks. Practitioners building upper-limb assistive devices could adopt the RL-trained musculoskeletal controller as a generalizable reference-torque generator and pair it with their own passivity-based safety layer for hardware deployment.

arXiv · cs.ROConceptual

Hijacking Robots with a Piece of Paper: A Systematic Study of Physical Prompt Injection in VLM-Controlled Robots

A sheet of paper with sneaky text can trick a robot into ignoring your orders.

Robots increasingly use AI vision-language models that read both what they see and what they're told, then decide what to do. This study shows that a piece of paper with cleverly written text — like a fake sign or a fake authority command — placed in the robot's camera view can hijack its reasoning, making it disobey or misbehave without anyone touching its code. The researchers tested this on a sorting robot with 20 different trick prompts across various scenes and instructions, running the AI 'brains' behind it (GPT-4o, Gemini) through over 5,600 trials. It matters because it exposes a whole new kind of security hole: attackers don't need to hack software, they just need to put the right words in the robot's line of sight.

Technical view

The paper systematizes physical (indirect) prompt injection against VLM-controlled robotic planners, defining a four-category taxonomy — indirect signage, task redefinition, authority impersonation, and conflict injection — and builds a 20-prompt benchmark tested across 3 scene layouts and 3 command formulations. They ran 5,670 trials against three frontier VLMs (GPT-4o, Gemini 2.5 Flash, and presumably a third) acting as sorting-task planners, measuring attack success as a function of scene, command specificity, and rule explicitness. This is directly usable as a red-teaming benchmark or evaluation harness for anyone deploying VLM planners in physically grounded robotic pipelines.

arXiv · cs.ROBuildable

Nonvisual Classification of Ground-Condition by Artificial Proprioception in an Amoeba-Inspired Autonomous Walking Robot

A blob-shaped walking robot 'feels' the ground with its feet instead of looking at it.

Most robots figure out rough terrain by using cameras, but cameras can be slow, power-hungry, or fooled by lighting. This robot instead uses a sense like our own — proprioception, the feeling of where your limbs are and how much force they're under — combined with an accelerometer and eight foot pressure sensors. A brain-inspired computing technique called reservoir computing sifts through the noisy, jittery sensor signals produced while the four-legged robot walks, and reliably tells flat ground from rough ground. Once it knows the terrain, the robot automatically switches its walking style on the fly, which matters for building robots that can navigate real-world environments without relying on vision.

Technical view

The system fuses a 3-axis accelerometer with 8 plantar pressure sensors and feeds the combined, motion-noisy signal stream into a reservoir computing (RC) classifier — an approach well-suited to temporal, high-dimensional sensor data without heavy training overhead. It achieves high-accuracy binary ground classification (flat vs. rough) despite large signal fluctuations from dynamic quadruped gait, and demonstrates closed-loop on-site gait switching conditioned on the classification output. The authors also report per-sensor contribution analysis, useful for anyone optimizing sensor placement/cost on legged platforms doing nonvisual terrain sensing.

arXiv · cs.ROBuildable

JoyAI-RA 0.5: Scaling Robot Manipulation Learning via Dual Action Alignment

An AI learns to control robot hands by watching human hands on video, not just robot data.

Teaching robots to manipulate objects is hard because there isn't much real robot data — but there's tons of video of humans doing everyday tasks. The problem is that human video, simulations, and real robot recordings don't line up neatly: they show different bodies, different cameras, and often no labeled 'actions' at all, so naively mixing them confuses the AI instead of helping it. This system, called JoyAI-RA, gets around that by inferring hidden 'action' patterns just from watching how scenes visually change over time, and separately by matching up trustworthy human and robot movements directly. Together these two tricks let the robot learn general physical common sense — how gravity and touch work — from cheap human videos and then transfer that skill to real manipulation tasks, which matters because it could make robot training dramatically cheaper and more scalable.

Technical view

JoyAI-RA 0.5 is a generalist Vision-Language-World-Action (VLWA) model that addresses negative transfer when pooling heterogeneous manipulation data (human egocentric video, simulation, real robot) via two alignment mechanisms: implicit action alignment, which infers latent actions from visual state transitions to let action-free data condition a learned world model of physical dynamics, and explicit alignment, which grounds reliable human/robot trajectories into a shared representation. The world-model-conditioning approach lets unlabeled video contribute dynamics priors even without ground-truth action labels, addressing the core embodiment-gap problem in cross-domain robot learning. This is relevant to practitioners building generalist manipulation policies who want to leverage large-scale egocentric video (e.g., Ego4D-style corpora) alongside limited teleoperated robot data.

arXiv · cs.ROBuildable

KILVO: Kinematic-Inertial-LiDAR-Visual Odometry with Robust Multimodal Adaptation for Humanoid Robots

A humanoid robot tracks its own position by combining leg feel, balance, laser scans, and camera vision.

For a walking humanoid robot to know exactly where it is and how it's moving, it needs to fuse information from several imperfect senses: an inertial sensor for balance, joint encoders that sense how the legs are moving, a laser scanner (LiDAR) that maps surroundings, and a camera. KILVO combines all four using a statistical filtering technique that predicts motion from the inertial and leg data first, then corrects that prediction using the laser and camera readings, run in a carefully staggered order to keep things fast and accurate. It's also built to degrade gracefully — if one sensor fails or gives bad readings, the system adapts rather than falling over, and it even estimates when and where the robot's feet are touching the ground. This matters because reliable self-localization is the backbone of any humanoid robot that has to walk safely through unpredictable real-world spaces.

Technical view

KILVO fuses joint encoders, IMU, LiDAR, and camera in an asynchronous-sequential hybrid error-state iterated Kalman filter (ESIKF): IMU drives the prediction step, leg kinematics are incorporated asynchronously at high rate as proprioceptive constraints, and exteroceptive updates are applied sequentially — first LiDAR point registration for geometric priors, then photometric visual updates. The design explicitly targets multimodal robustness/graceful degradation under partial sensor failure, and includes a compact contact estimation module that shares state information with the estimator, tackling the classic legged-robot foot-contact ambiguity problem. This is a concrete, implementable state-estimation architecture for anyone building odometry stacks on sensor-rich humanoid platforms.

arXiv · cs.ROBuildable

Search-Aided Joint Agent-Environment Reinforcement Learning for Robust Lifelong Multi-Agent Path Finding with Rotations

Warehouse robots learn to plan paths and even rotate in place without crashing into each other.

Warehouses full of robots constantly need new routes as they finish tasks and get assigned new ones, and existing AI planners often assume robots can move in unrealistically simple ways, ignoring things like needing to physically rotate in a tight space. This paper studies a more realistic version of the problem that adds safety margins and turning constraints, which makes coordinating many robots in tight spaces much harder. Their solution, SJRL, pairs learned robot behavior with a fast, one-step search algorithm that steps in to resolve collisions the learned policy misses and shares that fix back with nearby robots. This matters directly for real automated warehouses (think Amazon-style fulfillment centers), where efficient, collision-free, realistic robot movement saves time and prevents costly jams.

Technical view

The paper formalizes LMAPF-R2, a lifelong multi-agent pathfinding variant incorporating robust safety margins and in-place rotation constraints reflective of real warehouse robot kinematics, which substantially raises coordination difficulty in constrained spaces versus standard grid-based LMAPF. Their method, Search-Aided Joint Reinforcement Learning (SJRL), augments learned neural policies with Causal PIBT, a lightweight single-step search-based collision resolver that propagates conflict-avoidance decisions across agents rather than relying purely on end-to-end learned coordination. This hybrid learning+search approach is a practical template for anyone deploying RL-based MAPF policies who needs stronger safety guarantees than pure learned policies typically provide.

arXiv · cs.RORunnable

PathCover: A Fast Convex Decomposition along a Path via Randomized Iterative Space Partitioning (RISP) on Point Clouds

A new algorithm carves safe flight tunnels through cluttered 3D scans almost instantly.

When a robot plans a path through a cluttered space — say a drone flying through a room scanned by LiDAR — it helps to know which nearby chunks of empty space are safe to move through, described as simple bulging 'corridor' shapes called convex polytopes. Building these safe corridors quickly enough to keep up with a fast-moving sensor feed has been a bottleneck for existing methods. PathCover introduces a new randomized algorithm (RISP) that builds these safe shapes directly from raw point-cloud data in roughly the time it takes to just read through the data once, guaranteed to finish and to keep making progress along the intended route. This matters because faster, guaranteed safe-corridor generation directly speeds up and de-risks real-time robot trajectory planning, like drones or self-driving vehicles dodging obstacles.

Technical view

PathCover is a corridor-generation framework built on RISP (Randomized Iterative Space Partitioning), a novel algorithm that constructs convex polytopes directly from raw point clouds in expected linear time under a mild probabilistic elimination condition, rather than requiring costly preprocessing or global optimization. It produces sequences of overlapping obstacle-free polytopes suitable as safe-region constraints for downstream MPC or trajectory optimization, with a formal proof of finite-step termination and continuous progress along an arbitrary obstacle-free reference path. Benchmarks on synthetic and real LiDAR datasets show an order-of-magnitude speedup over prior state-of-the-art corridor generators, making it a strong drop-in replacement for real-time navigation pipelines that are compute-bound on corridor construction.

arXiv · cs.ROBuildable

ARGUS: Aligning Robot Scene Geometry Under Shifting Views with Large 3D Vision Models

AI teaches robots to see objects the same way no matter which camera angle is watching.

Robots trained to manipulate objects from camera images often accidentally learn 'where things appear in the picture' rather than 'where things actually are in space,' so if you move the camera, the robot gets confused even though nothing about the task changed. ARGUS fixes this by using powerful 3D vision AI models to mentally 'rotate' whatever the camera sees into one standard, canonical viewpoint before feeding it to the robot's decision-making system. This means the robot's policy always reasons about a consistent view of the world, regardless of where the physical camera happens to be mounted. It matters because it lets robots trained on varied datasets (from different labs, different camera setups) generalize far better to new viewpoints they've never explicitly seen.

Technical view

ARGUS is an observation-preprocessing pipeline that disentangles scene geometry from camera viewpoint by using large-scale 3D vision models to re-project arbitrary camera views into a canonical frame before passing observations into a downstream visuomotor policy, addressing the well-known viewpoint-overfitting failure mode in image-conditioned manipulation policies. It's evaluated across datasets with varying viewpoint diversity (echoing setups like DROID and BridgeV2), testing whether canonicalization improves generalization beyond training-distribution camera poses. As a modular preprocessing step, it's plug-compatible with existing visuomotor policy architectures, making it a low-friction way to boost cross-viewpoint robustness without retraining the policy backbone itself.

arXiv · cs.ROConceptual

Sliding Sensors: Configurable Confidence in State Estimation for Continuum Robots

A snake-like robot slides its sensor around to 'aim' its confidence wherever it matters most.

Continuum robots — flexible, snake- or tentacle-like machines — need to know their own exact shape and position to work safely, but sensors can only be placed in limited spots, so some parts of the robot are always known more precisely than others. The key insight here is that you don't always need the whole robot to be perfectly tracked — just the part doing the important work right now. This paper builds a physical prototype where a single sensor can slide back and forth along the robot's body, letting you dynamically choose where the robot is most 'sure' of itself. Sliding the sensor over time actually gives better overall shape estimates than leaving one sensor fixed in place, which matters for making flexible robots — useful in things like surgical tools — safer in tight, uncertain spaces.

Technical view

The work introduces a mechanically reconfigurable sensing concept for continuum robot state estimation, where a sensor can be longitudinally translated within the robot body to reshape the spatial confidence profile of the estimator rather than treating estimation uncertainty as spatially uniform. A hardware prototype demonstrates that estimation confidence at task-relevant locations can be actively tuned by relocating the sensor, and that periodically sliding the sensor over time reduces full-body shape-estimation error compared to a single fixed sensor position. This is an early-stage hardware/estimation concept most useful to researchers designing task-adaptive sensing strategies for soft/continuum robots (e.g., surgical or inspection manipulators) where uniform full-body accuracy is unnecessary or infeasible.

arXiv · cs.ROBuildable

World-to-Wrist: Task-Conditioned Future Wrist Modeling for Fine-Grained Robot Manipulation

Robots imagine how their own wrist camera view will change before deciding what to do next.

This is about robots that use cameras to pick up and manipulate objects — one camera watching the whole scene, another mounted right on the wrist for close-up work. Normally these two views are just fed into the AI side by side, which wastes the fact that the wrist view is special: it shows exactly what's about to happen during delicate contact, like fingers closing on a screw. The researchers built a system that first predicts how the wrist-camera view will evolve given the task, then uses that imagined near-future to decide the actual motor commands. Think of it like a mechanic mentally rehearsing the next few seconds of a fiddly repair before moving their hands — it helps with precise, fine-grained manipulation tasks that trip up ordinary robot-control AI.

Technical view

W2-VLA augments a vision-language-action (VLA) model by introducing latent 'modeling tokens' that bridge the VLM and a dedicated wrist-future predictor: conditioned on task instructions and wrist observation history, it forecasts future wrist-view latents which are then folded back into the context used for action decoding. This decouples global task semantics (main view) from local contact-level dynamics (wrist view) rather than treating them as parallel, undifferentiated inputs. A companion pipeline, W2-CoT, is introduced for synthesizing training supervision/chain-of-thought data for this future-modeling objective. Practitioners working on fine-grained manipulation (insertion, threading, tool-use) could adopt the latent-token interface as a plug-in module atop existing VLA backbones.

arXiv · cs.ROBuildable

Unified Planning-Learning Framework for Robust UUV Navigation Under Partial Observability

An underwater robot learns to map, plan, and dodge obstacles using only sonar, no GPS.

Autonomous underwater vehicles can't use GPS once submerged, and their sonar and depth cameras only reveal a fraction of their surroundings at any time — a classic 'partial observability' problem. This paper combines several techniques into one pipeline: the robot builds a memory-persistent map from what it's sensed so far, a planner charts a long-range safe route while keeping healthy clearance from obstacles, and a reinforcement-learning policy (trained by trial and error) handles split-second dodging of things nearby. It also compresses its noisy, incomplete sensor data into a compact internal 'belief' about the environment and adds a rulebook-like decision layer ('behavior tree') trained in stages for extra safety. The goal is more reliable underwater navigation for tasks like inspecting pipelines or mapping the seafloor where things are dynamic and visibility is poor.

Technical view

The framework fuses persistent occupancy mapping (built solely from onboard sonar/depth imagery), a clearance-constrained global planner for long-horizon structure, and an RL policy for short-range reactive tracking/avoidance, unified via a learned compact latent state that encodes environment structure, obstacle dynamics, and uncertainty under partial observability. A behavior-tree distillation step with staged supervision is used to improve training stability and safety guarantees, and an uncertainty-aware component (cut off in the abstract) further informs decision-making. This hybrid planning-learning architecture is relevant to anyone building navigation stacks for GPS-denied, sensor-limited robots — the occupancy-mapping-plus-RL-plus-BT-distillation combo is a reusable template beyond just UUVs.

arXiv · cs.CVBuildable

LoDA: A Level of Detection Aware Method and a Multimodal Sensing Benchmark for Object Level Change Detection

A new method spots exactly what changed in a city's 3D laser-scan map — and how confident it is.

Self-driving cars and smart-city systems rely on detailed 3D maps built from LiDAR laser scanners, but cities constantly change — buildings go up, trees get trimmed, cars move — so these maps need updating. Older methods for detecting such changes just compare two scans and flag differences point by point, without saying how big a change has to be before it 'counts' or grouping points into recognizable objects. This new pipeline instead identifies discrete objects in the scans, accounts for the fact that the scanner's own positioning has some error, and labels each detected change (like 'moved,' 'added,' 'removed') with a confidence score based on how it shifted in height, volume, or surface angle. It matters because reliable, object-level change detection keeps digital maps trustworthy for navigation and city planning instead of drowning users in noisy false alarms.

Technical view

The pipeline decouples registration, geometry, and semantics: detection-limit-aware point-cloud registration propagates pose/alignment uncertainty into spatially varying detection thresholds, geometry-driven object proxies are built via rule-based semantic and instance segmentation, and displacement cues (height, volume, surface-normal direction) are combined to assign one of five object-level change labels with an associated confidence score. This addresses the tile-based, threshold-driven limitation of prior raster-diff and point-cloud-network approaches that output ungrounded per-point scores. The paper also introduces a multimodal sensing benchmark, presumably enabling standardized comparison of object-level change-detection methods — useful for teams building HD-map maintenance pipelines who need calibrated confidence rather than raw anomaly scores.

arXiv · cs.ROBuildable

Failing Gracefully: Mitigating Impact of Inevitable Robot Failures

When a home robot inevitably breaks or glitches, this teaches it to fail safely, not catastrophically.

Robots that work around the house — near kids, pets, furniture — will eventually crash, malfunction, or misread a situation no matter how well they're engineered. Instead of only trying to prevent failures, this research asks a different question: when a failure does happen, how bad could it be, and for whom? The team built a way to estimate both the chance that a failure leads to some kind of harmful interaction (like bumping into a person) and how severe that interaction would be, so the robot can plan its actions to keep worst-case outcomes mild even while still getting its job done efficiently. They also built FailBench, a physics-simulated (MuJoCo) testbed so researchers can systematically stage robot failures and study the fallout in a safe, repeatable way, rather than testing on real hardware where mistakes are costly.

Technical view

The paper formalizes a safety metric combining (1) probability of impactful robot-entity interactions conditioned on failure occurrence and (2) severity of resulting outcomes, then uses this to inform planning decisions that trade off safety against task efficiency under a nonzero failure-probability assumption — a shift from failure-prevention to failure-impact-mitigation. FailBench, a MuJoCo-based simulation framework, provides a systematic environment for injecting/studying robot-environment failure interactions (software crashes, hardware degradation, unpredictable contact), presumably enabling reproducible benchmarking of safety-aware planners. This is directly applicable to service-robot planners that need graceful-degradation policies rather than binary safe/unsafe classifications, and FailBench looks reusable as an evaluation suite for other failure-aware control methods.

BIO

Biology

45 new
bioRxiv · genomicsConceptual★ flagship

Pangenome discovery and characterization of human protein-coding duplicated genes

Reading 298 full human genomes to find hidden, duplicated genes the reference missed.

Parts of our genome exist as near-identical duplicated blocks, and genes sitting in those blocks are notoriously hard to read — sequencing tools get confused by the repetition, so many such genes are missing or wrong in the standard human reference. Using long-read sequencing (which reads DNA in long continuous stretches, resolving the confusion) across 298 assembled human genomes plus billions of full-length RNA reads from 83 tissues, the researchers systematically map these tricky gene families. They discovered thousands of genes that vary in copy number between people and aren't in the reference at all, and showed many are genuinely active — carrying intact protein-coding instructions and switching on strongly in brain, embryo, and testis. They also corrected hundreds of gene models, including reclassifying some 'dead' pseudogenes as truly functional. This fills real gaps in our map of what human genes exist.

Technical view

Combining 298 long-read assembled human genomes with 5.6 billion full-length cDNA reads from 83 tissues, the authors phylogenetically interrogated 493 gene families in high-identity segmental duplications, discovering 2,713 potentially copy-number-polymorphic genes absent from the reference. For reference SD gene families with assignable paralog specificity, 60.0% are expressed with maintained open reading frames and 45.7% show high expression in brain, embryo, or testis. They revised 386 gene models, including 150 absent or divergent from T2T-CHM13 annotation and 236 (35.1%) pseudogenes reclassified as protein-coding on expression/ORF evidence. This provides a pangenome-scale, evidence-backed resource for SD gene annotation that groups building disease and evolutionary studies can integrate to recover previously invisible protein-coding genes.

bioRxiv · immunologyConceptual★ flagship

Lysosome-related organelle genes are required for mitochondrial transformations during bacterial infections in Caenorhabditis elegans.

Tiny worm cells reshape their power plants to fight bacterial infection — and need recycling organelles to do it.

When bacteria invade a cell, they feed on the cell's own mitochondria — the tiny power plants that store lipids, proteins, and iron the bacteria crave — so cells often rearrange their mitochondria as an early defense. But how an infection actually triggers that rearrangement has been a mystery. Using the roundworm C. elegans infected with dangerous bacteria like Staph aureus or Pseudomonas (or subjected to low oxygen), the researchers watched the mitochondrial network remodel itself. They then found that this remodeling depends on a set of genes controlling 'lysosome-related organelles' — specialized cellular recycling and storage compartments — and that those same organelles also help switch on the cell's broader infection-fighting genes. So a recycling system inside the cell turns out to be a key relay between sensing bacteria and mounting a mitochondrial and immune response.

Technical view

In C. elegans, infection by Staphylococcus aureus or Pseudomonas aeruginosa, or hypoxia, triggers remodeling of the host mitochondrial network. The study identifies lysosome-related organelle (LRO) genes as required for this infection-induced mitochondrial remodeling, and further shows LROs are needed to precipitate downstream infection-response gene expression. This mechanistically links a pathogen/hypoxia stimulus to mitochondrial morphological transformation via LRO function, implicating LROs as upstream mediators of both mitochondrial and transcriptional immune responses. Researchers can follow up by dissecting which LRO gene products and cargo (e.g., iron or lipid handling) drive the mitochondrial response and testing conservation of this LRO–mitochondria axis in mammalian innate immunity.

bioRxiv · cell biologyRunnable★ flagship

Spatial transcriptomics resolves ductal TREM2+ and stromal FOLR2+ macrophages in the normal human breast at microanatomical resolution

A microscopic map reveals two distinct macrophage 'security teams' patrolling different zones of healthy breast tissue.

Macrophages are immune cells that act like the body's cleanup and repair crew, and this study maps exactly where different types sit inside normal human breast tissue. The question is that these cells aren't all the same — they take on different jobs depending on their neighborhood — but nobody had pinned down how they're physically arranged in healthy breast. The researchers used tools that read which genes each individual cell is switching on while keeping track of its exact location in the tissue, so they can see cell type and address at once. They found one type (marked by a protein called TREM2) tucked right against the milk-duct walls, and another type (marked by FOLR2) spread through the connective tissue between structures. Knowing what 'normal' looks like matters because breast cancer hijacks these same cells, so this healthy baseline gives a reference for spotting when things go wrong.

Technical view

The authors integrate single-cell RNA-seq with Xenium in situ spatial transcriptomics to resolve macrophage transcriptional states at microanatomical resolution in normal human breast. They delineate a TREM2+ population intercalated between basal-myoepithelial cells (the human analog of mouse ductal-niche macrophages) and a FOLR2+ population distributed across interlobular stroma, with unbiased niche analysis further resolving a periepithelial FOLR2+ subpopulation. The pairing of dissociated scRNA-seq for depth with imaging-based in situ profiling for spatial context is the key methodological move, tying transcriptional identity to defined tissue niches. Practitioners can use these marker sets and niche definitions as a healthy-tissue reference to contrast against tumor-associated macrophage states, or extend the Xenium panel to probe ligand-receptor niche signaling.

bioRxiv · genomicsRunnable

Massively parallel characterization of adolescent idiopathic scoliosis risk variants

Scientists tested thousands of DNA snippets to find which ones actually cause scoliosis risk.

Adolescent idiopathic scoliosis — a sideways spinal curve that shows up in teens — is known to run partly in families, and genome-wide studies have flagged regions of DNA linked to it. But most of those regions are in 'non-coding' DNA (stretches that don't directly build proteins), so scientists didn't know which specific genetic letters actually do anything. Here, researchers used a technique that can test thousands of DNA sequences at once, in living cells, to see which ones ramp gene activity up or down — like testing thousands of light switches simultaneously to see which ones actually control a light. They ran this in cartilage cells (chondrocytes), since cartilage is thought to be central to how scoliosis develops, and pinpointed 92 specific variants that meaningfully change gene activity — a concrete shortlist of likely functional culprits rather than a vague genomic neighborhood.

Technical view

Using massively parallel reporter assays (MPRAs) in two human chondrocyte cell lines (TC28a2, SW1353), the authors tested 7,173 candidate regulatory sequences covering 1,664 variant positions in linkage disequilibrium with 26 GWAS-identified AIS lead variants, comparing 1,664 reference alleles against 4,708 alternate alleles for differential regulatory (enhancer/promoter-like) activity. This identified 92 variants with statistically significant allele-specific regulatory effects, providing a prioritized, functionally-validated candidate list for follow-up (e.g., CRISPR perturbation, eQTL colocalization, or mechanistic studies of chondrocyte gene regulation in AIS). The MPRA approach itself is a template replicable for functionally fine-mapping non-coding GWAS hits in other musculoskeletal or complex-trait loci.

bioRxiv · pharmacology and toxicologyBuildable

Schema-Grounded Multitask Instruction Fine-tuning for Joint Biomedical Named Entity Recognition and Relation Extraction in Pharmacovigilance

One AI model reads drug-safety papers and pulls out both the chemicals and how they're linked, at once.

Pharmacovigilance is the ongoing work of monitoring drugs for safety, which requires combing through mountains of scientific papers to extract facts like 'this chemical' and 'that disease are connected by this side effect.' Traditionally, software does this in two separate steps — first find the entities (drugs, diseases), then separately figure out how they relate — using different custom-built systems for each, which is inefficient and doesn't let the two steps learn from each other. This work instead fine-tunes one large language model with explicit instructions to do both jobs at once, guided by a structured 'schema' (a formal blueprint of what an entity or relationship should look like) across three different benchmark datasets. The payoff is a single, more scalable AI system that can pull structured, standardized safety information out of raw medical text with less custom engineering per task.

Technical view

The method jointly instruction-tunes an LLM for biomedical named entity recognition (NER) and relation extraction (RE) using a unified generative, schema-grounded multitask formulation, trained/evaluated across three benchmark corpora spanning chemical, disease, and drug entities plus chemical-disease and drug-adverse-event relations. This contrasts with the conventional pipeline of corpus-specific architectures trained separately for NER and RE, aiming for cross-task knowledge sharing and improved scalability. Practitioners building pharmacovigilance or biomedical literature-mining pipelines could adopt this schema-grounded instruction format as a template for adding new entity/relation types without retraining separate extractors, provided the schema is well specified.

bioRxiv · physiologyConceptual

FoxO factors preserve airway epithelial homeostasis by coordinating adaptive stress responses

A stress-sensing protein helps lung linings survive low oxygen, heat, and chemical damage — across three species.

The cells lining our airways are constantly bombarded by stressors — low oxygen, temperature swings, toxic chemicals — and need a way to sense trouble and mount the right defensive response. FoxO proteins act like an internal alarm-and-dispatch system: when stress hits, they move into the cell's nucleus (its control center) and switch on genes suited to that particular threat. The researchers tested this across three very different organisms — fruit flies, mice, and human cells — and found FoxO responds to hypoxia, heat, and oxidative stress, though different versions of the human FoxO protein specialize in different jobs. Because fruit flies have just one FoxO gene (unlike the several redundant copies in mammals, which make it hard to study), the team could cleanly show that removing it makes flies more vulnerable to low-oxygen stress on their airway lining specifically — pointing to FoxO as a fundamental, evolutionarily conserved guardian of airway health.

Technical view

The study shows that FoxO transcription factors undergo stress-induced nuclear translocation in airway epithelial cells (AECs) across Drosophila, mouse, and human, integrating extrinsic/intrinsic signals (hypoxia, temperature, oxidative stress) into transcriptional stress responses; human hFOXO paralogs show distinct, cell-type-specific activation patterns, indicating functional specialization/redundancy that complicates loss-of-function work in mammals. Using Drosophila's single dfoxo ortholog to bypass this redundancy, loss-of-function analysis showed hypoxia uniquely triggers a dfoxo-dependent innate immune response and that dfoxo is required for hypoxic stress resistance in the airway epithelium. This establishes FoxO as an evolutionarily conserved node for airway stress adaptation, offering a genetically tractable Drosophila model and candidate hFOXO paralogs for follow-up mechanistic or therapeutic studies of stress-related airway pathology (e.g., in hypoxia-associated lung disease).

bioRxiv · neuroscienceConceptual

Intertwined Autophagy and Integrin Dynamics Shape Axon Growth and Regeneration

Cells recycle their own machinery to help injured nerve fibers regrow after damage.

Autophagy is a cell's built-in recycling system — it breaks down and reuses worn-out internal parts to keep the cell healthy. Scientists already suspected autophagy plays some role in how nerve fibers (axons) grow and respond to injury, but exactly how wasn't clear. Separately, it's known that 'integrins' — proteins that anchor cells to their surroundings and are key to how axons physically extend — get moved around and recycled in a process that, in non-nerve cells, is regulated by autophagy. This study directly watched, in living adult sensory neurons under a microscope, how autophagy-related structures and integrin-carrying compartments behave inside axons both normally and after the axon is cut, to see whether that same autophagy-integrin partnership operates in neurons. Understanding this connection matters because it could reveal new ways to boost nerve regeneration after injury, relevant to conditions like nerve damage or spinal cord injury.

Technical view

Using live imaging in adult sensory neurons, the authors tracked autophagic vesicle dynamics and integrin-containing compartments within axons under basal growth conditions and following axotomy, testing whether autophagy-mediated regulation of focal adhesion turnover (previously shown in non-neuronal cells) extends to neuronal axon growth and regeneration. The abstract is cut off before stating the specific finding, but the setup establishes a live-imaging framework linking autophagosome trafficking to integrin-based focal adhesion dynamics as a candidate regulatory axis for axon extension and injury response. This positions autophagy-integrin coupling as a mechanistic target for researchers studying axon regeneration therapeutics, with live-imaging of vesicle/integrin co-trafficking as a reusable assay for follow-up perturbation studies (e.g., autophagy gene knockdown and its effect on regenerative axon outgrowth).

bioRxiv · neuroscienceConceptual

Pathway-specific short-term synaptic dynamics and lateral inhibition shape frequency-dependent input integration and population-level pattern separation in the dentate gyrus

Your brain's memory-sorting region uses timing tricks and inhibition to keep similar memories apart.

The dentate gyrus is a small brain region that takes overlapping, similar signals coming in from elsewhere in the brain and turns them into distinct, non-confusable patterns — a process called pattern separation that's key to telling apart similar memories (like where you parked today versus yesterday). This study builds a detailed computer model of the neurons and their three different input pathways, each of which strengthens or weakens differently depending on how fast signals arrive (their 'short-term synaptic dynamics'). By combining these pathway-specific dynamics with lateral inhibition — neighboring cells suppressing each other — the model shows how the circuit filters and separates information based on signal frequency. The payoff is a mechanistic explanation for how the brain avoids muddling similar memories together.

Technical view

The authors construct a three-tier computational modeling pipeline — from a 37-compartment biophysically detailed granule-cell model to network-level simulations — parameterized with Tsodyks-Markram short-term plasticity fits from slice electrophysiology for the lateral perforant path, medial perforant path, and proximal inputs. They show the distal lateral pathway facilitates (strengthens with repeated input) while other pathways show distinct dynamics, and that combining these pathway-specific kinetics with lateral inhibition produces frequency-dependent gating of dentate gyrus integration. This links single-synapse biophysics to population-level pattern separation, providing a mechanistic, testable framework connecting known short-term plasticity parameters to circuit-level memory discrimination performance and robustness.

bioRxiv · ecologyConceptual

Nutrient environments shape amino acid auxotrophy and cross-feeding

Some 'broken' bacteria can secretly make their own nutrients depending on what's around them.

Auxotrophy is when a microbe has lost the genetic ability to make some essential building block, like an amino acid, and normally must scavenge it from its environment or from neighboring microbes that share resources ('cross-feeding'). Scientists usually treat this dependency as fixed and permanent, but this study tests whether it actually shifts depending on what nutrients are available. Using engineered strains of two common bacteria (E. coli and B. subtilis) each missing a different amino acid-making gene, the researchers grew them alone and together across 40 different nutrient environments and measured how well they survived. They found that many 'broken' strains could still grow fine without the amino acid they supposedly needed, and that how much bacteria share resources with each other varies a lot depending on conditions — reshaping our understanding of how microbial communities form and depend on each other.

Technical view

The authors used matched panels of six amino-acid auxotrophic knockout strains each in E. coli and B. subtilis, assaying monoculture and pairwise coculture growth across 40 distinct carbon/nitrogen source combinations. They found auxotrophic phenotypes are not fixed but environment-contingent — strains lacking specific biosynthetic enzymes nonetheless grew in a substantial fraction of amino-acid-free conditions, implying alternative biosynthetic routes or nutrient substitution effects. Cross-feeding magnitude and directionality varied substantially across species and environments, suggesting that community assembly models assuming static metabolic dependencies should incorporate environmental context-dependence of auxotrophy for more accurate predictions of microbial interaction networks.

bioRxiv · evolutionary biologyConceptual

Excessive feeding induces development of a colony-like body plan in paratomous flatworms

Feed a tiny worm too much and it grows into a chain of connected clones instead of splitting apart.

Most colonial animals — like corals or coral-like sea creatures — are permanently built from repeated connected units called zooids, but scientists don't fully understand how solitary animals might have evolutionarily transitioned into that colonial lifestyle. This study looks at a microscopic flatworm, Stenostomum, that can either split into separate individuals (asexual fission) or stay connected in chains resembling a mini colony. The researchers discovered that simply feeding the worms more food shifts the balance between two growth processes — how fast the body elongates versus how fast a new head forms — and pushes the worm toward forming chains instead of splitting apart. This shows that something as simple as food supply can flip an animal between a solitary and colony-like body plan, offering a live model for how colonial life might have evolved.

Technical view

Using four Stenostomum species that facultatively alternate between asexual fission and tail-to-head zooid chains, the authors combined ecological manipulation (food availability), developmental and physiological measurements, and transcriptomics to show that excess feeding triggers chain (colony-like) formation. Mechanistically, increased nutrition alters the relative rates of somatic longitudinal growth versus head morphogenesis, biasing development toward incomplete individuation and multi-zooid chains rather than complete fission. This establishes a tractable, environmentally-triggered model system for studying the developmental and molecular basis of the solitary-to-colonial transition, complementing prior work in obligately colonial taxa like bryozoans and tunicates.

bioRxiv · geneticsConceptual

Ancient DNA reveals matrilineal organisation and recurrent unions between dominant matrilines in Iron Age Britain

Ancient DNA shows Iron Age Britons traced family lines through women, not men.

Kinship — who counts as family and how power and land pass down — shapes every traditional society, and archaeologists have long debated how much of this was really reflected in biology versus just social custom. Here, researchers sequenced genome-wide DNA from 534 people buried roughly 2,000+ years ago in Iron Age northeast England, mostly at a site called Wetwang Slack. By comparing genetic similarities across hundreds of skeletons, they reconstructed a massive 13-generation family tree of 195 individuals and found it was organized around the female line — meaning descent, identity, and likely social status tracked through mothers rather than fathers, unusual for known ancient societies. This is one of the strongest genetic proofs yet of a matrilineal society in prehistoric Europe, reshaping assumptions about gender and power in the ancient past.

Technical view

The authors generated genome-wide ancient DNA data for 534 individuals from Arras Culture sites (390 Wetwang Slack, 100 Pocklington, 29 Melton) in Middle Iron Age northeast England, using kinship inference and pedigree reconstruction methods. At Wetwang Slack they resolved a 13-generation pedigree of 195 individuals organized around matrilineal transmission, with recurrent reproductive unions occurring between members of dominant matrilines — evidence that social organization was structured by maternal descent lines rather than patrilineal or bilateral kinship. This is among the largest ancient DNA pedigree reconstructions to date and provides direct genetic evidence for matrilineal social structuring in prehistoric Britain, a rare finding relative to most documented ancient Eurasian societies.

bioRxiv · animal behavior and cognitionConceptual

Hoverfly responses to looming stimuli depend on elevation and speed

Hoverflies dodge things swooping in from below far better than from above.

When something is about to crash into you, it creates a rapidly expanding ('looming') image on your retina, and many animals have specialized brain cells and reflexes tuned to react instantly to this warning sign. In hoverflies, this matters for avoiding predators, obstacles, and even rival hoverflies. This study looked at the brain cells that detect looming threats and found they respond much more strongly to things looming from below (the ventral visual field) than from above. The researchers then tested real hoverflies, tethering them and showing looming shapes from different angles and speeds, to see if their evasive behavior actually matched this brain bias — confirming that both neural wiring and real escape behavior are tuned to threats from underneath, likely because that's where predators like birds are more dangerous.

Technical view

The authors characterize looming-sensitive descending neurons carrying visual threat signals from hoverfly optic lobes to thoracic motor centers, finding a strong bias toward sensitivity in the ventral visual field. They then tested tethered hoverflies with looming stimuli presented at different elevations (dorsal vs. ventral) and expansion speeds, measuring behavioral evasive responses to determine whether motor output matches the neural elevation bias. The results demonstrate elevation- and speed-dependent tuning of collision-avoidance behavior consistent with descending neuron physiology, offering a neuroethological model linking specific descending neuron populations to quantifiable escape behaviors, useful for comparative work on visual threat-detection circuits and bio-inspired collision-avoidance sensors.

bioRxiv · bioengineeringBuildable

Photoinitiator-free visible-light bioprinting: minimising radical-induced oxidative stress and DNA damage in gelatin hydrogels

A new 3D-printable gel gels under visible light without the toxic chemicals usually needed.

To 3D-print living tissue, scientists often use light to harden ('crosslink') a gel scaffold around cells, but this normally requires chemical additives called photoinitiators plus UV or harsh short-wavelength light, both of which can damage cells and DNA. This paper introduces a gel made from gelatin (a natural protein) chemically modified with light-reactive pyrene groups, which lets it harden directly under ordinary visible light with no added chemicals needed. The gel forms quickly, can be tuned to be softer or firmer, and stays structurally stable for over a month in cell culture — meaning it could let researchers print tissue scaffolds that are gentler on the cells growing inside them and safer to work with.

Technical view

The authors synthesize gelatin functionalized with acrylamidylpyrene groups (Gel-Pyr) that undergoes photoinitiator-free crosslinking via visible-light-induced [2+2] cycloaddition between pyrene moieties, eliminating the free-radical chemistry responsible for oxidative stress and DNA damage in conventional photoinitiator-based bioinks. Gel-Pyr shows rapid, tunable gelation kinetics, mechanically adjustable hydrogels, precise temporal control of crosslinking, and long-term (>30 day) structural stability in culture, characterized via rheology. This offers bioprinting practitioners a drop-in, radical-free alternative crosslinking chemistry compatible with standard visible-light bioprinting setups, potentially improving cell viability and genomic integrity in printed constructs.

bioRxiv · bioinformaticsBuildable

RVQ-Alpha: Bridging Single-Cell Transcriptomics and Large Language Models via Hierarchical Discrete Tokenization and Fact-Aware Reinforcement Learning

Teaching AI chatbots to 'read' individual cells like sentences in a language.

Single-cell RNA sequencing measures which genes are active in individual cells, producing continuous numerical data, while large language models like ChatGPT only understand discrete chunks of text called tokens — so there's been no clean way to let an LLM directly reason about cell biology. This paper introduces RVQ-Alpha, a system that converts each cell's gene activity pattern into a compact 'vocabulary' of discrete symbols the language model can natively process, plus a decoder that can translate those symbols back into the original gene expression data. It also trains the model with reinforcement learning that keeps its outputs grounded in real, named genes and facts rather than vague guesses. The goal is to let AI language models reason about and answer questions about cell biology as fluently as they do about text.

Technical view

RVQ-Alpha applies multi-codebook Residual Vector Quantization to lexicalize continuous single-cell expression profiles into a hierarchical discrete token alphabet embedded directly in an LLM's native token stream, paired with a decoder for expression reconstruction, addressing the representational mismatch between continuous scRNA-seq data and autoregressive token-based LLMs. The pipeline adapts four standard LLM training stages — tokenization, supervised fine-tuning, reinforcement learning, and distillation — to this domain, using 'Evidence-First' supervision to ground learned symbols in named genes and expression facts rather than unconstrained latent codes. This provides a concrete architecture for building gene-grounded, instruction-following LLMs over single-cell data, useful to practitioners building cell-state QA, annotation, or reasoning systems that need explicit gene-level interpretability.

bioRxiv · plant biologyRunnable

The N-Terminus of Sophora tonkinensis Cytochrome P450s Evolves Neutrally yet Encodes Rich Functional Information: A Protein Language Model Analysis

An AI language model finds hidden design rules in a floppy, fast-evolving piece of an enzyme.

Cytochrome P450 enzymes are proteins that perform key chemical reactions in cells, and each one is anchored to a membrane by a short chunk of its sequence at the very start (the N-terminus) — a piece that's essential for the protein to work but whose exact sequence varies wildly between species, making it hard to study with normal comparison methods. Here, researchers fed 345 versions of this enzyme from a medicinal plant into an AI protein language model (like ChatGPT, but trained on protein sequences) that can spot patterns without needing to line up matching sequences letter-by-letter. They found that even though the sequence looks like it's evolving almost randomly, it still contains hidden structural instructions — a tiny conserved 'hinge' plus two separate layers of encoded information, one about how the protein anchors itself in the membrane and another that distinguishes which enzyme family it belongs to. This shows AI models can uncover functional design rules invisible to traditional evolutionary analysis.

Technical view

The authors apply an alignment-free ESM2 protein language model pipeline to 345 Sophora tonkinensis cytochrome P450 sequences, showing the N-terminal 50-residue membrane anchor evolves under pervasive neutral relaxation (mean embedding dispersion 0.62) except for a highly conserved PxxG structural hinge (P21/G24, 98-99% conservation). Despite weak sequence-level constraint, decomposing the ESM2 embedding space reveals two separable information channels: an unsupervised method (DIVA) recovering a topological template for membrane-anchoring architecture — validated against physicochemical properties and DeepLoc-2.1 localization predictions — and a supervised ablation method (PIVOT) isolating a family-discriminating signal. This demonstrates that PLM embeddings can decompose functionally distinct constraints within a single low-conservation region, offering a template for alignment-free functional annotation of variable protein termini beyond classical sequence conservation analysis.

bioRxiv · neuroscienceConceptual

Shared texture-like representations, not global form, underlie deep neural network alignment with human visual processing

AI models predict brain activity not by "seeing objects" like us, but by matching textures.

Neural networks trained to recognize objects turn out to be surprisingly good at predicting how our visual brain responds to images, and scientists assumed this meant the networks had learned to 'see' objects the way we do. But this study suggests that's mostly wrong: even AI models with no training at all can predict brain signals, and getting better at object recognition doesn't make the match better. Instead, the researchers think both AI and our mid-level visual brain areas are keying in on texture-like patterns — the fine-grained visual 'stuff' (like grain, color, and pattern statistics) that fills objects and backgrounds — rather than the overall shape or identity of objects. They tested this by showing 57 people real photos and versions stripped down to just their texture statistics while recording brain activity (EEG). The finding reshapes what we think AI-brain similarity actually proves about how vision works.

Technical view

The authors dissociate object-related from texture-like representational content by recording 64-channel EEG from 57 participants viewing natural scenes alongside texture-synthesized counterparts that preserve summary statistics (à la Portilla-Simoncelli) while destroying global form. Using representational similarity analysis between EEG responses and DNN layer activations, they test whether DNN-brain predictivity tracks texture-statistic similarity rather than object-recognition accuracy, building on prior findings that untrained DNNs and accuracy-decoupled models still predict neural responses above chance. The result argues that reported DNN-brain 'alignment' in ventral visual cortex substantially reflects shared low/mid-level texture coding rather than shared object-recognition computations. Practitioners using DNN-neural predictivity as a benchmark for object recognition models should control for texture-statistic confounds, and this texture-synthesis + EEG paradigm offers a replicable template for isolating representational content driving such alignment.

bioRxiv · neuroscienceConceptual

Disruption of the Homer1 coiled-coiled domain by a novel de novo human HOMER1 variant impairs protein scaffolding, calcium signalling, and synaptogenesis

One swapped letter in a brain gene breaks how neurons build connections and sense calcium.

Homer1 is a protein that acts like scaffolding inside brain cells, holding other molecules in place so neurons can properly sense calcium, grow, and form connections with each other. This study found a new, spontaneously occurring mutation in the human HOMER1 gene (a single amino-acid swap called R297W) in a patient, and showed that the faulty protein doesn't just fail to work — it actively interferes with the normal, healthy copies too, like a broken cog jamming the whole machine. In lab-grown sensory neurons, cells carrying this mutant protein couldn't properly steer their growing tips toward a guidance chemical (BDNF), because the calcium-signaling system they depend on was disrupted. This matters because scaffolding-protein glitches like this are increasingly linked to conditions such as epilepsy and autism, so understanding exactly how one mutation derails neuron wiring helps explain those disorders' roots.

Technical view

The authors identify a de novo HOMER1 missense variant (R297W) that disrupts the protein's coiled-coil domain, the region mediating Homer1b/c tetramerization and scaffold assembly with mGluRs, IP3 receptors, and Ca2+ channels. Functionally, Homer1b/c^R297W acts in a dominant-negative manner: in dorsal root ganglion sensory neurons it impairs BDNF-gradient-guided growth cone turning and significantly blunts store-operated Ca2+ entry (SOCE), a process normally dependent on intact Homer1 scaffolding. This links a single structural lesion to downstream deficits in Ca2+ signaling, dendritic spine morphogenesis, and synaptic plasticity pathways implicated in epilepsy and ASD. Researchers could use this variant as a model system (e.g., in iPSC-derived neurons or DRG cultures) to dissect SOCE-dependent axon guidance mechanisms or to screen for compounds that rescue dominant-negative Homer1 scaffolding defects.

bioRxiv · neuroscienceConceptual

The Cellular and Synaptic Actions of Dopamine During Behavior

Scientists eavesdropped on live neurons to catch dopamine actually rewiring brain circuits in real time.

Dopamine is the brain chemical famous for driving reward and motivation, and it's long been known to be crucial for how we learn and move, but exactly how its brief, second-to-second signals translate into lasting changes in brain circuits has been a mystery. This is hard to study because you need to watch individual neurons' electrical activity while an animal is awake and behaving, and also control dopamine levels at the same time — a serious technical challenge. Here, researchers recorded the actual voltage changes inside single brain cells (in a region called the striatum) in mice while simultaneously boosting or suppressing their dopamine signals, watching what happened to the connections coming in from the cortex. Surprisingly, short-term dopamine changes (seconds to minutes) barely budged how strongly those connections worked, hinting that dopamine's big effects on behavior may build up differently than assumed. This matters for understanding conditions like Parkinson's and addiction, where dopamine signaling goes wrong.

Technical view

Using in vivo whole-cell patch-clamp recordings combined with real-time optogenetic/chemogenetic bidirectional manipulation of dopamine signaling in awake, behaving mice, the authors directly measure how dopamine dynamics shape corticostriatal synaptic transmission and intrinsic excitability at cell-type resolution. Contrary to models predicting rapid dopamine-dependent synaptic plasticity, acute (seconds-to-minutes) dopamine manipulations produced only modest changes in synaptic transmission and no detectable shifts in membrane potential dynamics or intrinsic excitability — implying that behaviorally relevant plasticity may require longer timescales or different induction mechanisms than commonly assumed. This whole-cell-plus-dopamine-manipulation approach is a methodological advance practitioners can adopt to directly test dopamine-dependent plasticity rules in specific striatal cell types (e.g., D1 vs D2 MSNs) rather than inferring them from indirect behavioral or imaging proxies.

bioRxiv · ecologyConceptual

Global patterns of helminths associated with gelatinous zooplankton: a missing link in marine parasite transmission

Jellyfish-like blobs are secretly major highways for parasitic worms across the world's oceans.

Gelatinous zooplankton — jellyfish, comb jellies, and their relatives — are everywhere in the ocean, both eating and being eaten, yet nobody had really mapped out their role in spreading parasitic worms (helminths) through marine food webs. The researchers pulled together nearly 450 records of these worm-jelly relationships from dozens of studies worldwide, and added their own new discoveries from five seas, using both microscopy and DNA analysis to identify the parasites. They found a surprising pattern: unlike most ocean life, which is most diverse near the equator, these worm infections were actually most common and diverse in temperate (mid-latitude) waters, and the Red Sea had far more infected jellies than the nearby Mediterranean, while none turned up in the Baltic or North Seas. This matters because it reveals gelatinous zooplankton as an overlooked 'missing link' that ferries parasites between smaller prey and larger predators like fish, reshaping how scientists think about disease transmission in ocean ecosystems.

Technical view

The study synthesizes 431 host-parasite association records (from 89 published sources plus 23 new records) documenting helminths using gelatinous zooplankton as intermediate or paratenic hosts, combining morphological identification with molecular (likely barcoding/phylogenetic) confirmation from Mediterranean, Red, Celtic, Baltic, and North Sea samples. Contrary to the classical latitudinal diversity gradient (higher parasite diversity near the tropics), helminth occurrence and richness peaked at temperate latitudes, with striking regional heterogeneity (high prevalence in the Red Sea, zero detections in Baltic/North Seas). This positions gelatinous zooplankton as a significant, previously underappreciated trophic conduit in marine parasite life cycles; researchers building on this could use the compiled dataset to model transmission pathways to higher trophic levels or test drivers (salinity, temperature, host density) behind the observed latitudinal reversal.

bioRxiv · evolutionary biologyConceptual

Modulation of Avian Iridescence via Melanogenesis

The same pigment genes that color feathers also secretly build peacocks' shimmering rainbow nanostructures.

Peacock feathers and many other birds' iridescent colors don't come from pigments but from microscopic structures in the feathers that bend light like a prism — physicists have understood this optical trick for a while, but nobody knew what genes actually build those tiny structures. This study looked at peacocks bred in captivity that had lost or changed their iridescence, plus wild birds whose feathers are iridescent on one side and dull on the other, and used genetic and single-cell analysis to find the genes responsible. Surprisingly, all the genes involved turned out to be ones normally known for controlling melanin, the pigment that makes hair and skin dark, not genes specifically for building nanostructures. This shows that the pigment-making machinery does double duty, also assembling the physical scaffolding needed for structural color, connecting two things — color from pigment and color from physics — that scientists usually treat as separate.

Technical view

Combining genomic analysis of domesticated Indian peafowl mutants with single-cell transcriptomics of naturally occurring asymmetric feathers in wild bird species (iridescent vs. non-iridescent barbules on the same feather), the authors identify eight genes underlying gain, loss, and shifts of iridescent structural color in peafowl. All causal peafowl mutations mapped to canonical melanogenesis genes, indicating that melanosome biogenesis/organization machinery is co-opted to assemble the periodic nanostructures (melanosome arrays) that produce iridescence via thin-film/multilayer interference. This provides a genetic and cellular mechanism linking pigmentary and structural coloration pathways, giving researchers specific candidate genes to manipulate (e.g., via CRISPR in avian models or comparative genomics across iridescent lineages) to test causal roles in nanostructure assembly and photonic tuning.

bioRxiv · bioinformaticsRunnable

jazzPanda: spatially aware marker gene detection for imaging-based spatial transcriptomics

A new software tool finds which genes mark cell types by using their actual location, not just counts.

When scientists use spatial transcriptomics — techniques that show exactly where each gene is active inside a tissue sample — they still need to figure out which genes serve as reliable 'name tags' for each cell type. The problem is that the standard tools for finding these marker genes were built for older technology that doesn't track spatial position, and they struggle when transcript counts per cell are very sparse, which is common with imaging-based spatial methods. jazzPanda solves this by converting the raw locations of cells and gene transcripts into simplified spatial 'signal' vectors through a binning process, and can even work directly from transcript coordinates without needing to painstakingly outline individual cells first. It then flags marker genes using statistical methods (like correlation tests or a regularized regression model) that account for one or multiple samples. This gives researchers a more reliable, spatially-aware way to identify what makes each region or cell type of a tissue distinct.

Technical view

jazzPanda is a computational method for marker gene detection tailored to imaging-based spatial transcriptomics (e.g., MERFISH/Xenium-style data), addressing the sparsity and spatial-blindness limitations of scRNA-seq-derived marker detection tools. It aggregates cell and transcript coordinates into 1D vectors via spatial binning, enabling marker identification directly from raw transcript coordinates without requiring cell segmentation — useful when segmentation is unreliable or unavailable. Markers are called using either permutation-based rank correlation for single-sample datasets or a lasso-regularized generalized linear model for multi-sample/multi-condition designs. Practitioners working with sparse imaging-based spatial transcriptomics data can adopt jazzPanda as a drop-in analysis step to identify spatially informative markers where standard scRNA-seq marker methods (e.g., Seurat's FindMarkers) underperform.

bioRxiv · plant biologyConceptual

GSNOR-dependent nitric oxide homeostasis promotes recovery from repeated climate stress across generations in Arabidopsis thaliana

A single gene lets plants 'remember' climate stress and recover better across five generations.

As climate change brings more droughts, heat, and pollution, scientists want to know whether plants can somehow pass on resilience — or damage — from one generation to the next, and what controls that. This study grew a small mustard-family plant (Arabidopsis) for five generations, exposing some to repeated stresses like drought, high CO2, ozone, and heat, while comparing normal plants to a mutant missing a gene called GSNOR that helps regulate nitric oxide, a signaling molecule cells use to respond to stress. After three stressed generations, they let two generations recover under normal conditions and tracked plant growth, photosynthesis, seed output, and which genes were switched on or off. The mutant plants missing GSNOR consistently grew worse and produced fewer seeds than normal plants across environments, showing that properly managing nitric oxide levels is important not just for surviving stress in the moment, but for a plant lineage's ability to bounce back generations later.

Technical view

Using a five-generation experimental evolution design (G1-G3 under control, drought, elevated CO2, O3, warming, and combined-stress treatments; G4-G5 recovery under control conditions), the authors compare Arabidopsis thaliana wild-type (Col-0) to the gsnor1-3 knockout, which disrupts S-nitrosoglutathione reductase-dependent nitric oxide (NO) homeostasis. They quantify rosette growth, photosynthetic traits, seed yield, and transcriptome dynamics (RNA-seq) across all generations and scenarios, finding gsnor-ko shows consistently reduced vegetative growth and reproductive output relative to WT, with strongly scenario- and generation-dependent transcriptomic shifts. This establishes GSNOR-mediated NO homeostasis as a mechanistic node in transgenerational stress memory/recovery, giving researchers a genetic entry point (gsnor mutants) and a transcriptomic dataset to dissect which NO-responsive pathways mediate multigenerational resilience versus stress inheritance in plants.

bioRxiv · plant biologyBuildable

Pathogen-dependent biocontrol activity of Chlorella sorokoniana aqueous extracts against fungal and oomycete plant pathogens

Algae water extract stops a devastating rice-killing fungus almost as well as some pesticides.

Fungal and mold-like pathogens destroy huge amounts of crops every year, and farmers need alternatives to chemical pesticides that are gentler on the environment. This study tested whether a simple water-based extract from Chlorella sorokoniana — a common microalgae — could fight off three major crop-damaging fungi and oomycetes (mold-like pathogens), using lab dishes, detached leaves, and whole living plants. The extract worked dramatically well against one particular fungus, Magnaporthe oryzae, which causes rice blast disease: it cut fungal growth by about 70%, blocked the fungus's ability to punch into leaf cells by 55%, reduced leaf damage by over 75%, and lowered overall disease severity by nearly 65% when applied preventively. However, its effectiveness varied a lot depending on which pathogen it was tested against, showing that this algae-based biocontrol isn't a one-size-fits-all solution but could be a targeted, more sustainable tool for specific crop diseases like rice blast.

Technical view

The study screens aqueous extracts of the microalga Chlorella sorokoniana against three economically significant fungal/oomycete phytopathogens using a tiered in vitro (mycelial growth inhibition), ex vivo (detached leaf lesion assays), and in planta (whole-plant disease severity) pipeline. Efficacy was strongly pathogen-dependent, with the strongest activity against Magnaporthe oryzae (rice blast): ~70% in vitro growth inhibition, 55% suppression of appressorium formation (the specialized infection structure fungi use to penetrate host cells), >75% reduction in detached-leaf lesion development, and 64.5% reduction in disease severity with preventive application in planta. This establishes C. sorokoniana extract as a candidate biocontrol agent specifically effective against appressorium-dependent pathogens like M. oryzae; researchers could next fractionate the extract to identify the active antifungal compound(s) and test formulation/field efficacy or mechanism of appressorium inhibition.

bioRxiv · plant biologyBuildable

The chloroplast CLP chaperone-protease system controls the steady-state abundance of the singlet oxygen sensors EXCUTER 1 and 2

Plants use a protein-shredding machine to decide how many stress-alarm sensors survive in their solar panels.

Chloroplasts, the solar panels of plant cells, contain a quality-control machine called the CLP protease that chops up and recycles damaged or unwanted proteins. Two 'adaptor' proteins, CLPS1 and CLPF, help this machine recognize which proteins to destroy. Researchers used a genetic trick — a modified CLPC1 chaperone that 'traps' proteins right before they're destroyed — in normal plants and in plants missing one or both adaptors, to see which proteins depend on each adaptor. They found this same destruction machinery also controls levels of EXECUTER1, a protein that senses a toxic form of oxygen made when chloroplasts are stressed, showing how plants tune their own stress-alarm system by degrading the alarm-sensor itself.

Technical view

Using an in vivo CLPC1 substrate-trapping approach (35S:CLPC1-TRAP-STREPII) in Arabidopsis WT, clpf, clps1, and clpfclps1 backgrounds, the authors mapped adaptor-dependent CLP substrate recruitment. CLPF trapping was reduced in clps1, consistent with a physical CLPS1–CLPF interaction bridging substrates to CLPC1. EX1, the chloroplastic singlet-oxygen sensor, was efficiently trapped in WT and clps1 but not in clpf, indicating CLPF is the primary adaptor mediating EX1 turnover, while clpfclps1 lethality upon TRAP expression underscores non-redundant, essential adaptor functions. This positions CLP-mediated proteolysis as a regulatory node controlling ROS-sensor abundance and retrograde singlet-oxygen signaling.

bioRxiv · neuroscienceConceptual

Departure from OFF-State Microstate Dynamics Tracks Levodopa Response in Parkinson's Disease

Brain-wave 'snapshots' reveal how Parkinson's drugs nudge a frozen brain back into motion.

Parkinson's disease depletes dopamine, a brain chemical needed for smooth movement, and doctors want to know exactly how dopamine-replacement drugs like levodopa change brain activity, not just muscles. Researchers recorded whole-brain magnetic signals (MEG) from 13 patients while off their medication and again about an hour after taking it, then broke the constantly shifting brain activity into brief repeating patterns called 'microstates,' like freeze-frames of a movie. They checked whether medication reliably shifted the mix of these patterns away from the 'off' pattern, and whether patients whose brain patterns changed more also moved better. This links an objective brain signal to how well treatment is actually working, which could eventually help tailor dosing.

Technical view

The study performed source-reconstructed MEG microstate analysis in 13 bradykinetic-dominant PD patients scanned OFF and ON levodopa, using microstate transition matrices to characterize whole-brain resting-state dynamics rather than static spectral power. The core analysis tests whether the magnitude of departure from the OFF-state microstate transition profile correlates with motor improvement, providing a dynamics-based biomarker of dopaminergic response. This positions microstate transition metrics as a candidate objective, network-level correlate of treatment efficacy, replicable with any source-localized MEG/EEG pipeline supporting microstate segmentation and transition analysis.

bioRxiv · synthetic biologyBuildable

PANCS-Inhibitors: A rapid method to directly select for protein-protein interaction inhibitors

A phage-powered 'matchmaker breaker' evolves molecules whose only job is prying apart cancer-driving protein pairs.

Many diseases, including cancers, are driven by two proteins sticking together when they shouldn't, and finding a drug to pry them apart usually starts by finding something that merely binds one protein and hoping it blocks the interaction — a slow, hit-or-miss process. This new method, PANCS-Inhibitors, uses bacteria-infecting viruses (phages) in a continuously evolving selection system, but instead of just selecting molecules that bind a target, it specifically selects molecules that actively break apart a pre-formed pair of proteins. Tested on three cancer-relevant pairs, including the well-known KRas-Raf and Mdm2-p53 pairings, it both improved existing blockers and discovered brand-new mini-protein blockers that work inside mammalian cells. This directness could sharply cut the time and luck needed to develop targeted drugs.

Technical view

PANCS-Inhibitors adapts Phage-Assisted Non-Continuous Selection to directly enrich for binders that disrupt a pre-formed PPI complex, rather than binders selected independently and screened post hoc for inhibitory activity, by coupling phage propagation/infectivity to disruption of a reporter-linked PPI. Validated against KRas-Raf, Mdm2-p53, and Myc-Max, the platform both optimized known inhibitors and identified de novo mini-protein inhibitors with confirmed activity in mammalian cells. Practitioners could adapt the selection circuit to other disease-relevant PPI targets for rapid inhibitor discovery without a separate binding-then-functional-screening cascade.

bioRxiv · molecular biologyConceptual

Structural basis for alternative 3' splice site selection in the human spliceosome active center

A newly spotted molecular glue-point helps cells pick the correct cut site when editing genetic messages.

When cells read genes, they must cut extra pieces from the raw genetic message (RNA) and splice the remaining parts together — but sometimes there are two nearby possible cut points, and choosing the wrong one causes disease-linked errors like those seen in BRCA1 and CFTR (the cystic fibrosis gene). Using cryo-EM, an imaging method that flash-freezes molecules to photograph their 3D shape, researchers discovered a previously unknown protein, SDE2, that stabilizes the spliceosome — the cell's molecular scissors and glue — right as it commits to a cut site, working alongside two other factors to favor weaker, less obvious cut sites. When this stabilizing role is disrupted, faulty splicing worsens, and restoring it partly fixes the RNA errors seen in BRCA1 and CFTR. This gives a structural explanation for how cells normally avoid a class of splicing mistakes tied to real disease.

Technical view

Cryo-EM structures of the human spliceosome C* complex reveal SDE2 as a factor that stabilizes the catalytic active center specifically to enforce docking of weaker, proximal 3' splice sites, with a truncated SDE2(ΔN) construct trapping a stalled C* intermediate showing impaired docking-factor engagement. SDE2 works combinatorially with Prp18 and FAM32A, reading a cis-regulatory sequence code around competing 3'-ss to bias site selection. SDE2ΔN destabilizes 3'-ss docking in vitro and rescues BRCA1/CFTR mis-splicing events in vivo, directly implicating this docking-factor network in pathogenic AG-gain mutation outcomes and suggesting it as a target for correcting specific mis-splicing defects.

bioRxiv · molecular biologyBuildable

Aggrecan hypomorphism accelerates the progression of post-traumatic osteoarthritis in mice

Cartilage built with too little cushioning protein breaks down faster after a knee injury, in mice.

Cartilage, the cushioning tissue in joints, relies heavily on a molecule called aggrecan to stay springy and absorb shock, and after injuries, cartilage often degrades over time into osteoarthritis. Researchers used a special mouse line that lets them dial down aggrecan specifically in cartilage, then simulated a knee injury by surgically destabilizing a knee ligament in mice with normal, partial, or greatly reduced aggrecan levels. They tracked joint damage over three months using tissue scoring and micro-CT scanning, a detailed 3D X-ray, to see how fast arthritis progressed. Mice with less aggrecan to start developed worse, faster-progressing arthritis after injury, supporting the idea that baseline cartilage 'cushioning reserve' affects how vulnerable a joint is to injury-triggered arthritis.

Technical view

Using tamoxifen-inducible Agc1CreERT2 mice to generate graded aggrecan hypomorphism (heterozygous vs. homozygous), the authors performed DMM surgery to induce post-traumatic OA and assessed progression at 4, 8, and 12 weeks via OARSI, synovitis, and osteophyte scoring plus micro-CT, alongside sGAG quantification and assessment of matrix-degrading proteases and aggrecan/collagen II degradation neoepitopes. Reduced baseline aggrecan dose-dependently accelerated PT-OA severity and structural joint damage post-DMM. This establishes a genetically tunable hypomorphic model for dissecting how baseline proteoglycan reserve modulates OA susceptibility, useful for testing protease inhibitors or matrix-restorative therapies.

bioRxiv · cell biologyConceptual

Coronary Artery Disease Transcriptomics Reveals Two Drivers of the Endothelial Cell SR-BI Expression and LDL Transport that Underlie Atherosclerosis

Two separate biological triggers ramp up a cholesterol doorway in artery walls, feeding heart disease.

Heart attacks and strokes often start when LDL cholesterol, the 'bad' kind, leaks from the blood into artery walls, and a doorway protein called SR-BI on blood vessel lining cells controls how much gets in. Using single-cell RNA sequencing, a technique that reads which genes are active in each individual cell, on real diseased human arteries, researchers found this doorway protein is turned up in cells sitting where blood flow is disturbed or turbulent, like branch points prone to plaque buildup. In mice, they showed high cholesterol and turbulent blood flow independently crank up SR-BI, and traced the flow-driven increase to a stress-response protein called HIF-1 acting on a specific DNA switch in the SR-BI gene. Understanding these two separate triggers gives potential new targets for blocking the earliest step of artery-clogging disease.

Technical view

Single-cell RNA-seq of human CAD coronary specimens shows endothelial SR-BI (Scarb1) upregulation concentrated in atheroma-associated cells bearing a disturbed-flow transcriptional signature. In mice, hypercholesterolemia and disturbed flow independently upregulate endothelial SR-BI, with flow-driven upregulation specifically initiating endothelial LDL uptake and atherogenesis; transcription-factor network analysis identifies HIF-1 binding within Scarb1 intron 1 as the mechanism driving flow/hypercholesterolemia-associated SR-BI induction in vivo. This nominates the HIF-1–Scarb1 intron 1 axis as a druggable node for blocking endothelial LDL transcytosis specifically at atheroprone, disturbed-flow sites.

bioRxiv · cell biologyConceptual

Atypical RanGAP drives nucleocytoplasmic transport in a parasitic Alveolate

A parasite rewired a completely different protein to do a job every other cell relies on a specific gene for.

Every cell needs to shuttle large molecules between its nucleus, where DNA lives, and the rest of the cell, and this traffic is controlled by a protein called Ran, which needs a helper enzyme called RanGAP — a helper found in nearly every complex organism. Strangely, Toxoplasma gondii, a common parasite, appears to lack the usual RanGAP gene, so scientists purified proteins from the parasite to hunt for whatever was secretly doing that job. They discovered a repurposed protein, TBC9, normally used for a different cellular task, has evolved a new function as RanGAP, and showed it can substitute for the real thing in yeast and is essential for the parasite's nuclear traffic. This reveals an unexpected way essential cellular machinery can be rebuilt from spare parts, and since TBC9 is unlike the human version, it could be a weak point unique to the parasite.

Technical view

Through biochemical purification of RanGAP activity from Toxoplasma gondii lysates, the authors identify TBC9, a RabGAP-fold protein, as a neofunctionalized substitute for the canonical RanGAP that is absent in many alveolates/apicomplexans. TBC9 complements RanGAP function in yeast and is essential for nucleocytoplasmic transport in Toxoplasma, and in vitro assays with purified Toxoplasma Ran and TBC9 demonstrate specific, robust GAP activity despite the divergent fold. This establishes convergent evolution of RanGAP function from a non-canonical GTPase-activating domain, offering a structurally distinct, potentially selectively targetable node in apicomplexan parasites.

bioRxiv · cell biologyBuildable

Extracellular matrix particle treatment induces digit regeneration in soft-tissue preserved amputation (SPA) model of adult mice

Sprinkling ground-up tissue scaffold on an amputated mouse toe coaxes new bone to grow back.

Some animals can regrow amputated fingertip-like structures, but mice normally can't regenerate the middle bone of a toe once it's cut off. Researchers tried boosting the healing stump using extracellular matrix (ECM), the natural scaffolding material that surrounds cells and helps guide tissue repair, but as solid particles rather than liquid, which is hard to keep in place with standard amputation surgery. So they invented a new setup that removes the amputated bone but keeps the surrounding soft tissue intact, letting them pack ECM particles inside like a pouch. Using this new model, they found the ECM particles genuinely triggered new bone to form at the cut end, confirmed with detailed 3D micro-CT scans, suggesting a practical way to promote limb-tip regeneration after injury.

Technical view

The authors developed a soft-tissue preserved amputation (SPA) model in mice, removing the amputated P2 phalanx bone while retaining surrounding soft tissue, to enable retention of solid ECM particles at the amputation site, circumventing containment limitations of ECM solutions in classical amputation models. ECM particles implanted and wrapped within preserved soft tissue induced new bone formation at the distal P2 stump, quantified via morphological and micro-CT assessment. This establishes SPA as a tractable platform for testing particulate biomaterials in mammalian digit-tip regeneration, with direct applicability to optimizing ECM particle composition and dose for bone-regenerative therapies.

bioRxiv · cell biologyConceptual

Cell-type-specific decoding of Hippo pathway inactivation drives seminiferous epithelial collapse and rete testis hyperplasia in the adult mouse testis

Flip off one growth-control switch in mouse testis cells and sperm factories collapse while a nearby duct balloons.

Every tissue in the body needs signals telling cells when to grow and when to stop — the Hippo pathway is one of the master switches for that. Here scientists genetically deleted two genes (Lats1 and Lats2) that normally keep this brake engaged, doing so specifically in two testis cell types: Sertoli cells (which nurse developing sperm) and the cells lining the rete testis (a duct system that collects sperm). Removing the brake caused the sperm-producing tissue to rapidly degenerate, wiping out nearly all germ cells, while the rete testis duct grew abnormally large — showing that the same 'grow more' signal can push neighboring tissues in opposite directions depending on their identity. This matters because it reveals how fertility-critical testis architecture depends on precisely tuned growth signals, with implications for understanding infertility and tissue-specific cancer risk.

Technical view

Using WT1-Cre-driven conditional knockout of Lats1/Lats2 (core Hippo pathway kinases) in adult mouse testis, the authors ablated Hippo signaling in both Sertoli cells and rete testis epithelium simultaneously. The result was rapid seminiferous epithelial collapse with near-total germ cell loss, paired with marked rete testis hyperplasia, vascular remodeling, macrophage infiltration, and ECM deposition. Cell-type-resolved transcriptomics showed both lineages activate a shared core of YAP/TAZ target genes upon Lats1/2 loss, but their broader transcriptional programs diverge according to baseline lineage identity — a useful dataset for dissecting context-dependent Hippo/YAP output. Researchers building on this could mine the transcriptomic data to identify lineage-specific cofactors that redirect YAP/TAZ activity toward degeneration versus hyperplasia.

bioRxiv · cell biologyConceptual

Spatiotemporal transcriptomic landscape of synovial joint repair - an in vivo murine multimodal model of osteochondral injury

Mice with a small joint injury reveal, cell by cell, why cartilage repair usually goes wrong.

When cartilage and bone in a joint get injured — say from a sports mishap — the body's repair job is often sloppy, leaving the joint prone to arthritis later. This study created a small, controlled injury in mouse knee joints and then used a technique called spatial transcriptomics, which reads out which genes are switched on in each cell while also recording exactly where that cell sits in the tissue, combined with imaging and standard tissue staining over time. By mapping the whole joint's cellular response as it unfolds, the researchers aimed to catch the earliest, most decisive moments of repair — the ones that seem to set the tissue on a path toward healing or toward long-term damage. Understanding this timeline could point to when and how to intervene to help joints heal properly instead of degenerating into osteoarthritis.

Technical view

The authors established a reproducible non-critical osteochondral injury model in the trochlear groove of female C57BL/6 mice and profiled it with single-cell resolution spatial transcriptomics across the entire joint, layered with longitudinal micro-imaging and quantitative histology/immunophenotyping. The goal is to build a spatiotemporal atlas of the earliest cellular and molecular events in osteochondral repair — cell types, their locations, and their transcriptional states as repair progresses — which prior bulk or non-spatial approaches couldn't resolve. This gives regenerative medicine researchers a reference map to identify candidate cell populations or signaling windows for therapeutic intervention aimed at preventing post-traumatic osteoarthritis. The dataset and imaging pipeline could be reused to benchmark interventions against the model's natural (often incomplete) repair trajectory.

bioRxiv · developmental biologyBuildable

Shining stars: Transgenesis and efficient metamorphosis in the sea star Patiria miniata

Scientists gave a humble sea star glowing, gene-edited cells that survive all the way to adulthood.

The bat star, a common sea star used to study how animals develop, has long lacked the genetic tools that let scientists label and track specific cells — tools routine in mice or fruit flies. This team used CRISPR, a precise gene-editing technique, to insert a fluorescent tag directly into a gene the star uses everywhere in its body, making cells glow so they can be tracked under a microscope. They also found a genetic 'on switch' (a promoter) that can drive any inserted gene to turn on, and showed it keeps working even as the larva transforms into a juvenile star, a dramatic process called metamorphosis. They further worked out a fast, reliable way to trigger that metamorphosis in the lab, which is essential for anyone trying to raise transgenic sea stars to adulthood. This opens the sea star up as a genetically trackable model for studying development, regeneration, and reproduction.

Technical view

The authors used CRISPR/Cas9 to perform endogenous knock-in tagging of a ubiquitously expressed actin locus in Patiria miniata, both fusing fluorescent markers directly to the gene and isolating its promoter to drive standalone transgene cassettes delivered via plasmid. Critically, the transgene expression persists through metamorphosis into the juvenile stage, and the paper reports a fast, efficient protocol for inducing healthy larva-to-juvenile metamorphosis — previously a bottleneck for establishing stable transgenic lines. This toolkit (validated promoter + knock-in strategy + metamorphosis protocol) gives echinoderm researchers a practical path to CRISPR knock-ins, reporter lines, and stable transgenics in a classic developmental/regenerative biology model.

bioRxiv · developmental biologyBuildable

Dynamic Functional Pathway Development in Type 1 Spinal Interneurons: Stage-specific roles of retinoic acid activity

A vitamin-A-derived signal times exactly when spinal nerve cells divide, move, and finally mature.

As an embryo builds its spinal cord, nerve cells (interneurons) have to divide, differentiate into their final form, and physically migrate to the right spot — all in careful sequence. This study tracked one specific type of interneuron in quail embryos using single-cell RNA sequencing, a method that reads out the active genes in thousands of individual cells at once, letting researchers reconstruct a timeline of how a cell changes from a young dividing progenitor into a mature neuron. They compared normal embryos to ones deprived of retinoic acid, a vitamin-A-derived molecule known to guide development, and found it plays different jobs at different stages — first affecting cell division and positioning signals, later cytoskeleton and connective-tissue genes, and finally genes needed for the neuron to wire into circuits. This kind of staged molecular choreography helps explain how the nervous system reliably assembles itself, which matters for both basic biology and disorders of neural development.

Technical view

Using single-cell RNA-seq of the E4 quail neural tube under control versus retinoic acid (RA)-deprived conditions, the authors built a continuous pseudotime/RNA velocity trajectory of dI1 interneuron development from dorsal progenitors to differentiated neurons, cross-validated against in vivo medial-to-lateral cell displacement. They identify sequential transcriptional modules — early cell-cycle/BMP programs, a middle wave of cytoskeletal/ECM remodeling genes, and late pan-neuronal/synaptic genes — with RA loss disrupting stage-specific transitions, including altered BMP antagonist upregulation and apoptosis timing. The dataset provides a reference trajectory and RA-perturbation contrast that others can use to test candidate regulators of the proliferation-to-migration-to-differentiation handoff in spinal interneuron subtypes.

bioRxiv · developmental biologyConceptual

Screen Reveals Novel Roles for Tau, and Morphogen Gradients in Resolving an Epithelial Identity Crisis

A fly-wing screen catches a famous Alzheimer's protein moonlighting as a tissue bouncer.

Skin and other linings of the body (epithelia) constantly replace their cells, and mistakes during that turnover can seed cancers called carcinomas — so the body has quality-control systems that spot and remove misfit cells. Fruit fly wings are a classic testbed for this because their epithelium is simple and well studied; cells there carry molecular 'ID badges' called selector genes that mark where they belong, and cells with the wrong badge get eliminated. This paper describes a large, multi-stage genetic screen — systematically switching genes on or off and watching what breaks — to find new genes involved in keeping this quality control working. Among the surprises, they found roles for Tau, a protein best known for its involvement in Alzheimer's disease, and for gradients of morphogens, chemical signals that tell cells their position in the body. The findings connect fly wing biology to broader questions about how tissues police cell identity and prevent cancer-prone errors.

Technical view

The authors performed a three-tiered genetic screen in the bilayered Drosophila wing epithelium to identify novel regulators of epithelial homeostasis and selector-gene-based cell elimination (the process by which cells expressing incorrect positional/selector identity are removed from the tissue). Among the hits, they report previously unappreciated roles for Tau and for morphogen gradient signaling in resolving these 'epithelial identity crises.' This gives Drosophila epithelial biologists a new candidate gene list and screening pipeline to dissect the mechanics of cell-fitness surveillance, with potential relevance to mammalian epithelial quality control and carcinoma suppression given Tau's involvement.

bioRxiv · ecologyConceptual

Do microbial effects on hosts vary across life stage and vitalrate?

A massive meta-analysis finds microbes usually help their hosts consistently, not just at some life stages.

Many organisms, from plants to animals, live in partnership with microbes that can help or hurt them — theory predicts these effects might flip from good to bad depending on the host's life stage, since young and old (or growing versus reproducing) individuals may have very different needs. The researchers combed through published studies to compile 1,736 measurements of how microbial partners affected different aspects of host success — survival, growth, reproduction — across different life stages and species. Contrary to the expected mixed picture, they found that microbial effects usually point in the same beneficial direction across life stages rather than flipping between helpful and harmful. This suggests that many microbial partnerships offer benefits that hold up consistently over an organism's whole life, which is useful for predicting when symbiosis will be reliably good for a host rather than situational.

Technical view

This is a meta-analysis synthesizing 1,736 effect-size measurements of host-symbiont interactions across taxa, testing whether the sign (positive/negative) of microbial effects on host vital rates (survival, growth, reproduction) varies by host life stage, as predicted by demographic/evolutionary resource-allocation theory. Against that prediction, effects were found to be directionally consistent across life stages more often than expected under a null model, suggesting symbiotic benefits tend to be structurally stable rather than trade-off-driven across ontogeny. This provides a quantitative baseline and dataset for testing life-history theories of symbiosis, and researchers could stratify the same dataset by taxon or symbiont type to look for exceptions to the general consistency pattern.

bioRxiv · ecologyConceptual

Atmospheric nitrogen deposition and anthropogenic land use linked to changing fungal endophyte prevalence in cool-season grasses

128 years of herbarium seeds reveal how pollution and land-use change reshaped a hidden grass-fungus partnership.

Many wild grasses carry fungi living invisibly inside their seeds and tissues, called endophytes, which can boost the plant's ability to handle stress. As humans have farmed, paved, and polluted the landscape, it's unclear how these partnerships have fared over time. Rather than growing new experiments, this team turned to herbaria — collections of preserved, dated plant specimens — examining 8,739 seeds from nearly 2,000 specimens spanning 1895 to 2019 across three grass species, testing each for the presence of these seed-transmitted fungi. By linking historical fungal prevalence to records of nitrogen pollution and land-use change over the same period, they can see how human activity may have made this natural symbiosis more or less common. This kind of historical detective work matters because it shows whether long-standing plant-microbe partnerships are eroding under modern environmental pressures, with consequences for grassland resilience.

Technical view

The authors screened 8,739 seeds from 1,951 herbarium specimens of three grass species (Agrostis hyemalis, A. perennans, Elymus virginicus) collected 1895–2019 for presence of seed-transmitted Epichloë fungal endophytes, then correlated historic endophyte prevalence trends against records of atmospheric nitrogen deposition and anthropogenic land-use change. The approach exploits herbarium specimens as a longitudinal, century-spanning natural archive of symbiosis prevalence — a method applicable to other seed- or tissue-transmitted symbioses where fresh field sampling can't reach back in time. This provides a reusable framework and dataset for linking global-change drivers to shifts in plant-microbe symbiosis frequency at decadal-to-century scale.

bioRxiv · ecologyConceptual

Snow leopard-mediated apparent competition between wild ungulates and domestic livestock in the Trans-Himalaya

More wild prey near snow leopards means more, not fewer, livestock killed — a 13-year Himalayan puzzle.

Conservationists generally want more wild prey animals around for predators like the endangered snow leopard, partly hoping it reduces the incentive for these cats to kill livestock. But ecological theory says it could go either way: abundant wild prey might satisfy predators and spare livestock, or it might simply support more predators overall, increasing total livestock losses — a dynamic called apparent competition, where two prey types indirectly harm each other by sharing a predator. Using a rare 13-year dataset from a Himalayan region in India where snow leopards, wild ungulates, and herders' livestock coexist, researchers tested which pattern actually holds. They found that when wild prey numbers were higher, livestock depredation was also higher, not lower — supporting the apparent-competition scenario. This matters directly for conservation policy, since it complicates the assumption that boosting wild prey alone will protect herders' livestock and ease human-predator conflict.

Technical view

Using a 13-year time series from the Upper Spiti Landscape (Indian Trans-Himalaya), the authors tested whether wild ungulate density predicts snow leopard (Panthera uncia) depredation rates on free-ranging domestic livestock, distinguishing between apparent facilitation (wild prey buffering livestock losses) and apparent competition (shared predation pressure increasing losses) hypotheses. They found a significant positive association between wild ungulate density and livestock depredation that remained robust after bootstrapping to account for density-estimate uncertainty, supporting predator-mediated apparent competition rather than facilitation. This long-term carnivore dataset — rare given how elusive snow leopards are — offers a template for modeling multi-prey predator dynamics and informs livestock-compensation and prey-augmentation strategies in human-carnivore coexistence landscapes.

bioRxiv · biochemistryBuildable

GPMAW Glyco-Search: An Integrated Workflow for Identification and Validation of Intact Sialylated N-Glycopeptides

A smarter way to catch the sugary tags on proteins that keep slipping past mass spectrometers.

Many proteins are decorated with chains of sugar molecules called glycans, and a special acidic type called sialylation is especially hard to detect because it's rare, wildly variable in shape, and tends to fall apart during the standard lab technique (mass spectrometry) used to identify it. This paper builds a step-by-step pipeline: first chemically fishing out the sugar-tagged protein fragments, then running them through the mass spectrometer twice (once whole, once with sugars stripped off), and finally using new software called GPMAW to match the pieces back together with high confidence. It matters because these sugar tags influence how the immune system, cancer cells, and antibody drugs behave, so reading them accurately is key to biology and medicine.

Technical view

The workflow combines selective TiO2 enrichment for sialylated glycopeptides with dual LC-MS/MS runs on both intact and enzymatically deglycosylated glycopeptides. GPMAW's glyco-search uses the experimentally determined deglycopeptide backbone identity to constrain candidate glycan compositions before matching against intact-glycopeptide precursor masses, rather than searching glycan and peptide space jointly like conventional engines. Identifications are cross-validated using diagnostic oxonium ions, glycopeptide Y-ion fragment ladders, and an empirical glycopeptide confidence score, improving specificity for low-abundance, structurally heterogeneous sialoglycopeptides.

bioRxiv · biochemistryConceptual

Two activation heat capacity regimes underlie temperature-dependent catalysis in homologous archaeal ADP-dependent kinases

Why enzymes from icy, mild, and ancient microbes hit their speed limit at different temperatures.

Enzymes are the molecular machines that speed up chemical reactions in cells, and they usually work faster as things heat up—until a point where they suddenly slow down, even before the enzyme itself falls apart. A theory called MMRT explains this by saying the fleeting 'in-between' state during the reaction becomes unusually rigid at higher heat, and this paper asks whether that rigidity effect is the same across related enzymes adapted to very different temperatures. The researchers compared three versions of a related kinase enzyme: one from a cold-loving microbe, one from a middle-of-the-road microbe, and one reconstructed to resemble their shared ancient ancestor. They found the enzymes fall into two distinct behavioral patterns rather than one universal rule, suggesting evolution can tune this heat-sensitivity trait to match an organism's environment.

Technical view

The authors measured glucokinase activity of three homologous bifunctional ADP-dependent PFK/GK enzymes—MbPFK/GK (psychrotolerant, from Methanococcoides burtonii), MmPFK/GK (mesophilic, from Methanococcus maripaludis), and ancM (an ancestrally reconstructed variant)—across a temperature range and fit the data using macromolecular rate theory (MMRT) to extract activation heat capacity (ΔCp‡). Rather than a single conserved ΔCp‡ value, the enzymes cluster into two distinct heat-capacity regimes, indicating this parameter is evolutionarily tunable rather than a fixed universal catalytic signature. This supports incorporating lineage- and niche-specific ΔCp‡ measurements when using MMRT to predict or engineer temperature-adapted enzyme kinetics.

bioRxiv · bioengineeringBuildable

Evolution-inspired multi-objective Bayesian optimization for protein engineering

An AI that breeds proteins like evolution, but juggling several traits at once.

Protein engineering means searching for a sequence of amino acids that gives a protein some useful property, but the space of possible sequences is astronomically large and testing each one in the lab is slow and costly. EvoMOBO is a computer method that mimics evolution: it generates new protein variants step-by-step building on promising ones, has them 'compete' against each other, and uses a statistical technique (Bayesian optimization) to smartly pick which variants are worth testing next—while balancing multiple desired properties at the same time, like stability and activity together. Tested on real protein datasets, it out-performed existing methods at finding good, diverse candidates. This matters because it could shrink the number of expensive lab experiments needed to engineer better enzymes, antibodies, or other useful proteins.

Technical view

EvoMOBO is an active-learning framework combining path-dependent sequence generation, population-level competition among generated variants, and explicit multi-objective Bayesian optimization to navigate protein fitness landscapes under limited evaluation budgets. Benchmarked against state-of-the-art methods on the steroid receptor DNA-binding domain and ParD3 antitoxin landscapes, it showed robust enrichment toward target regions, advancement of the Pareto front, and maintained sequence diversity across two- and three-objective optimization tasks. Notably, in the DBD landscape it used simulation-derived geometric descriptors as labels for initialization and iterative updates, enriching for favorable measured activities without requiring experimental labels—a strategy practitioners could adapt when experimental fitness labels are scarce.

bioRxiv · bioengineeringBuildable

An integrated human forebrain organoid reveals microglia-mediated CD8⁺ T cell recruitment and neuroimmune dysfunction in Alzheimer's disease pathology

Lab-grown mini-brains with immune cells show how the body's own defenders may worsen Alzheimer's.

Organoids are small clumps of tissue grown in a dish from stem cells that mimic parts of a real organ—here, a miniature human forebrain. Researchers added two kinds of immune cells: microglia, the brain's resident cleanup crew, and T cells, immune cells that normally patrol the bloodstream, to see how they interact with Alzheimer's-like damage. They found microglia do double duty—clearing away the sticky amyloid-beta protein clumps linked to Alzheimer's and helping neurons mature, but also triggering inflammation that calls in T cells, creating a vicious cycle of brain inflammation. Blocking specific chemical signals (receptors called CCR5 and CXCR3) stopped the T cells from being recruited, pointing to a possible new drug strategy for calming this neuroimmune spiral.

Technical view

The authors built a modular iPSC-derived forebrain organoid platform co-integrating microglia and CD8+ T cells to model innate and adaptive neuroimmune interactions in Alzheimer's pathology. Microglia mediate amyloid-beta clearance and neuronal maturation but also drive inflammatory activation and CD8+ T cell recruitment via CCL4/5 signaling through CCR1/5 and CXCR3 receptors, establishing a self-reinforcing neuroinflammatory feedback loop. Pharmacological blockade of CCR5 or CXCR3 abolished T cell recruitment and altered autophagy in a microglia-dependent manner, identifying these receptors as candidate targets and providing a reusable human-relevant platform for testing immunomodulatory AD therapeutics.

bioRxiv · bioengineeringBuildable

A biofilm-derived peptide as an underwater adhesive

Cholera bacteria's slime yields a peptide glue that sticks underwater.

Making adhesives that actually work underwater is hard, since water weakens most glues—nature's best examples come from mussels and barnacles. This study instead looks at bacterial biofilms, the slimy protective films bacteria build, specifically from Vibrio cholerae (the bacterium that causes cholera), and pulls out a small protein fragment called a peptide from that slime. Using microscopes, computer simulations, and mechanical pull tests, they show this peptide sticks firmly to different wet surfaces and can even clump tiny particles together like a flocculant (a substance used to gather fine particles out of liquid). They also managed to manufacture the peptide cheaply using E. coli, a common lab bacterium, hinting it could be scaled up for real-world adhesive or water-treatment use.

Technical view

The authors characterize a peptide derived from Vibrio cholerae biofilms as a candidate underwater adhesive, using confocal microscopy, molecular dynamics simulations, atomic force microscopy, lap shear testing, and spectroscopy to assess surface adsorption and wet-bonding strength across substrates. The peptide also functions as a flocculant, co-aggregating with microspheres, and was successfully recombinantly expressed and purified from E. coli, demonstrating scalable production. This establishes bacterial biofilm proteins as a novel, genetically engineerable design space for underwater adhesives, complementary to existing mussel-foot-protein-inspired chemistries.

bioRxiv · bioengineeringRunnable

Characterizing the Assembly and Functional Properties of Gene-Length Mixed DNA Monolayers on Electrodes for Cell-Free Expression

Genes wired straight onto gold chips that print glowing proteins on command.

Cell-free protein expression means making proteins using just the molecular machinery from cells, without needing living cells themselves—useful for biosensors and lab-on-a-chip devices. Here, researchers stuck entire genes directly onto gold electrode surfaces using a strong chemical bond (thiol-gold chemistry), creating a reusable template that can churn out a glowing test protein (GFP) right where it's attached. They tested how tightly genes could be packed on the surface, and how factors like applied electrical voltage, storage time, repeated use, and protein buildup affected whether the chip kept working. The results show these gene-covered electrodes can be reused under some conditions but lose activity under others, informing how to design durable bioelectronic devices.

Technical view

Thiol-modified sfGFP genes were assembled into gene-length DNA monolayers on planar gold electrodes via thiol-gold chemisorption, with surface density tunable through DNA incubation concentration; immobilized genes support direct cell-free transcription-translation of sfGFP from the electrode surface. The authors systematically probe monolayer stability and reusable expression output against applied voltage, storage, repeated reaction cycles, reducing agents, and protein fouling, finding that dense chemisorbed monolayers retain partial function under some but not all tested conditions. This dataset is directly useful for engineers designing reusable bioelectronic or cell-free synthetic biology platforms that need stable DNA templates on conductive substrates.

Q

Quanta — Explained

1 new
Quanta MagazineConceptual★ flagship

Neutrinos From Deep Inside Earth Provide a New Picture of the Mantle

Ghostly particles streaming out of Earth's interior let us 'see' the radioactive heat driving the planet.

Neutrinos are nearly weightless particles that pass through solid matter almost undisturbed, and radioactive elements deep inside Earth constantly emit them as they decay. Scientists want to know how much of Earth's internal heat — the energy that moves continents and drives volcanoes — comes from this radioactive decay versus leftover heat from the planet's formation. Because these 'geoneutrinos' fly straight out from wherever the decay happens, a worldwide network of huge underground detectors can catch a rare few and work backward to estimate how much uranium and thorium sits in the mantle. This is essentially using particle physics as a telescope pointed inward, giving a picture of the deep Earth we can't reach by drilling. It matters because it tests our fundamental models of how the planet is built and how it stays geologically alive.

Technical view

Geoneutrino detection exploits inverse beta decay signals in large liquid-scintillator detectors (e.g., KamLAND, Borexino, and increasingly JUNO) to count antineutrinos from U-238 and Th-232 decay chains in the crust and mantle. Combining a global constellation of detectors at different crustal settings lets researchers subtract the modeled crustal contribution and constrain the mantle's radiogenic heat budget, feeding directly into bulk-silicate-Earth composition models and the total radiogenic-vs-primordial heat partition. The core claim is that multi-site data are now precise enough to meaningfully bound mantle heat production. Practitioners can build on this by folding new detector exposure and improved reactor-background and crustal-flux models into Bayesian inversions for mantle U/Th abundances.

HN

What's Trending

56 new
Hacker News · 1050 ptsConceptual★ flagship

What happens if an entire class of workers loses faith in their careers

When a whole profession stops believing its work has a future, the fallout ripples far beyond individual burnout.

This piece asks what happens when it's not just one burned-out person but an entire category of workers who collectively lose confidence that their careers still have a future. The real-world trigger is usually a big disruption — automation, industry decline, or technology that seems to make a skill obsolete — that makes people question whether their training and experience still count. The exploration is less about hard data and more about the human and economic consequences: how morale, recruitment, and the willingness to invest years in learning a craft all shift when faith collapses across a field. It matters because careers are built on the belief that effort compounds over time, and when a group stops believing that, the effects spread into hiring, mentorship, and the supply of skilled people society depends on. (The item is a title/opinion prompt with no supplied detail, so the specifics of its argument aren't given here.)

Technical view

This is a commentary/opinion item rather than a research artifact, so there is no method or dataset to characterize; it poses a labor-economics and organizational-psychology question about collective career pessimism within an occupational class. The relevant framing draws on concepts like occupational identity, expectancy of returns to human-capital investment, and cohort-level morale effects on labor supply and mobility. Absent the full text, the concrete claims and evidence can't be summarized without fabrication. A practitioner interested in the question could operationalize it via longitudinal surveys of career-confidence measures against entry/exit rates and reskilling uptake in an affected sector.

Hacker News · 868 ptsConceptual★ flagship

“Code was never the hard part” is an insult to all programmers

A pushback against the fashionable claim that writing code is the easy, throwaway part of software.

There's a popular saying — amplified by the rise of AI coding tools — that 'code was never the hard part,' implying the real work is design, product thinking, or coordination, while typing out code is trivial. This article argues that framing is dismissive and wrong, insulting to the craft and skill that actually writing good code demands. The core point is that translating a fuzzy idea into precise, correct, maintainable instructions is itself deeply hard intellectual work, not a mechanical afterthought. It matters right now because if teams and tools assume coding is 'solved,' they may undervalue the engineers and the careful judgment that keep software working. (This is an opinion piece stated as a title, so the detailed arguments aren't provided here.)

Technical view

This is an opinion/essay item, not a technical paper, responding to the widespread 'code is the easy part' narrative that has intensified with LLM code-generation tools. The implicit thesis is that the difficulty in software lies substantially within the act of encoding requirements into correct, precise, and maintainable code — spanning edge cases, invariants, and long-term maintainability — not solely in upstream design or product decisions. Without the article body, its specific supporting arguments cannot be enumerated without invention. Practitioners can engage with the debate by examining empirical measures of where defect cost and rework actually concentrate across the software lifecycle rather than accepting the rhetorical dichotomy.

Hacker News · 802 ptsConceptual

New Mexico court orders Meta to pay $567m over harms to children’s mental health

A US court just made Meta pay $567 million over harm to kids' mental health.

A court in New Mexico ruled against Meta, the company behind Facebook and Instagram, ordering it to pay $567 million into a fund aimed at addressing mental health harms tied to underage users of its platforms. The ruling is part of a broader wave of lawsuits and regulatory pressure across the US targeting social media companies over claims that features like addictive feeds and notifications damage teenagers' wellbeing. It signals growing legal consequences for tech platforms accused of prioritizing engagement over the safety of younger users.

Technical view

A New Mexico court judgment orders Meta to pay $567 million into a fund earmarked for addressing mental health harms to minors, part of ongoing state-level litigation (alongside multidistrict litigation and other state attorney general suits) alleging that Instagram/Facebook engagement-driving design choices—algorithmic feeds, notification patterns, and lack of adequate safeguards—contributed to compulsive use and psychological harm among underage users. The ruling adds to a growing body of legal precedent that could push platforms toward mandated design changes for minor-facing products.

Hacker News · 782 ptsRunnable

DeepSeek V4 Flash 0731

DeepSeek appears to have quietly shipped a new, faster AI model.

DeepSeek is a Chinese AI lab known for releasing capable large language models at unusually low cost, and this entry points to what looks like a new dated release, 'V4 Flash' from July 31, likely a faster and cheaper variant within its V4 model family. 'Flash'-style versions in AI naming conventions typically trade a bit of raw capability for speed and lower cost, aimed at high-volume everyday use rather than the most demanding tasks. Beyond the name and date, no further details are available here, so specifics like its exact capabilities or pricing would need to be checked directly from DeepSeek's own announcement.

Technical view

This is a title-only reference to "DeepSeek V4 Flash 0731", presumably a dated snapshot release of a faster, lower-cost "Flash"-tier variant within DeepSeek's V4 model family, continuing the industry pattern of frontier labs shipping distilled or lighter-weight variants alongside flagship models. No architecture, benchmark, context-length, or pricing details are given in the source, so any integration work should start by pulling the official model card/API docs from DeepSeek to confirm specifics before building against it.

Hacker News · 621 ptsConceptual

Danish high schoolers will have to verbally defend written assignments

Denmark will make teens defend their essays out loud to prove they actually wrote them.

Danish high schools are rolling out a new requirement: after turning in a written assignment, students will have to sit down and verbally explain and defend their work to a teacher. The problem is that AI chatbots can now write convincing essays in seconds, so teachers can no longer be sure a student actually understands — or even wrote — what's on the page. The fix is old-fashioned: an oral, exam-style conversation where the student has to justify their arguments and sources on the spot, which is much harder to fake with a tool that isn't in the room. It matters because it's one of the first concrete school-system responses to AI cheating that tests understanding directly instead of relying on unreliable AI-text detectors.

Technical view

The policy shifts assessment from artifact-based grading (the written document) to a hybrid model pairing the written submission with a mandatory oral defense, similar to a thesis viva. This sidesteps the unreliability of AI-text detectors by making comprehension and provenance verification part of the grading protocol itself. It requires curriculum and staffing changes, since individual oral sessions are far more labor-intensive per student than reading an essay, raising open questions about scalability across a full student body.

Hacker News · 556 ptsConceptual

Mea Culpa – Dark Hours

A security team publicly owns up to a mistake made while tangling with a crew called Dark Hours.

This appears to be an incident writeup or blog post where a security researcher or company admits — "mea culpa" is Latin for "my fault" — to an error connected to a threat actor or operation known as Dark Hours. Without more detail, the gist is that someone in the security world is being transparent about a slip-up, whether a misconfiguration, a bad call during an investigation, or flawed analysis, rather than quietly burying it. Public postmortems like this matter because the security industry runs on shared trust, and openly admitting mistakes helps others avoid repeating them.

Technical view

The title suggests a retrospective disclosure tied to an entity or campaign named Dark Hours, likely a ransomware group or threat-actor handle, where the author acknowledges an error in prior analysis, response, or attribution. Without the full text, the specific technical failure (e.g., misattribution, a flawed indicator of compromise, or an operational lapse) can't be confirmed. Readers tracking ransomware or threat-intel accuracy should treat it as a case study in correcting the record and improving verification practices in incident response.

Hacker News · 533 ptsConceptual

Oracle bans AI-generated code from OpenJDK

Oracle just told OpenJDK contributors: no code written by AI, full stop.

OpenJDK is the open-source project behind Java, one of the world's most widely used programming languages, and Oracle stewards it. Oracle has now banned contributors from submitting code generated by AI tools like ChatGPT or Copilot. The concern involves things like unclear copyright ownership of AI output, potential licensing contamination from models trained on other people's code, and worries about subtle bugs slipping past human reviewers. This matters because it's a major, high-profile project drawing a hard line against AI-assisted contributions just as AI coding tools become ubiquitous, and it could set a precedent other big open-source projects follow.

Technical view

The policy prohibits patches or code produced by generative AI tools from being submitted to the OpenJDK reference implementation, likely citing concerns over copyright provenance, license compatibility with the GPL+Classpath Exception, and the review burden of verifying AI-generated code quality. This mirrors similar debates in other major FOSS projects, such as some Linux kernel maintainers' skepticism toward AI-generated patches, and sets a governance precedent for large, legally sensitive codebases. Contributors will need to disclose or avoid AI tooling in their workflow, giving maintainers explicit grounds to reject suspect submissions.

Hacker News · 508 ptsConceptual

2027 memory capacity is reportedly sold out

Chipmakers say every byte of memory they'll make in 2027 is already spoken for.

"Memory" here means the RAM and storage chips, like DRAM, that go into everything from phones to data-center servers. Demand has exploded, largely driven by the AI boom, since building massive AI systems requires huge amounts of fast memory to train and run them. According to this report, manufacturers' entire production capacity for 2027 has already been booked by customers — a sold-out sign two years in advance. This matters because it signals a looming shortage: prices are likely to rise, and any company that hasn't locked in supply, including makers of ordinary laptops and phones, could face higher costs or delays.

Technical view

The claim is that DRAM and memory fabrication capacity for calendar year 2027 has already been fully allocated via advance purchase agreements, reportedly driven by hyperscalers and AI accelerator makers securing HBM (high-bandwidth memory) and conventional DRAM supply for AI training and inference clusters. This mirrors prior semiconductor supply cycles where multi-year fab capex lead times can't keep pace with sudden demand spikes. Expect elevated DRAM and NAND spot and contract pricing through 2026-2027, tighter allocation for non-hyperscaler buyers, and knock-on effects for consumer electronics costs.

Hacker News · 494 ptsRunnable

Fastmail offers EU data region

Fastmail now lets you keep your inbox stored entirely within EU borders.

Fastmail is an email provider, and it has just launched an option to host your inbox specifically in European Union data centers rather than wherever the company's servers happen to sit by default. This matters for privacy-conscious users and businesses because EU data protection law, GDPR, has strict rules about where personal data can legally live, and some organizations are required, or simply prefer, to keep data physically within EU jurisdiction to avoid it being subject to foreign government access laws like the US CLOUD Act. It's a straightforward but meaningful move: the same email service, now with a guarantee about the geography of your data.

Technical view

Fastmail is introducing an EU data residency option, letting customers pin their mailbox storage and processing to EU-based infrastructure rather than the provider's default hosting region. This addresses GDPR data-residency requirements and reduces exposure to extraterritorial legal demands, since data physically resident in the EU falls under EU jurisdiction for access requests. Organizations with compliance needs can provision new accounts into the EU region, likely as a selectable setting at signup or via a support-assisted migration for existing accounts.

Hacker News · 487 ptsBuildable

My server is a phone now

An old smartphone gets a second life as a tiny always-on home server.

This is a hobbyist project about repurposing an old, unused smartphone to run as a home server: instead of tossing a phone that still has a working battery, chip, storage, and network connection, someone installed server software on it to host things like websites, file storage, or home automation tools. The appeal is that phones are cheap, often free since people have old ones lying around, power-efficient, and come with a built-in battery backup and even a cellular connection — features a normal server or Raspberry Pi doesn't have out of the box. It's a fun, sustainable example of reducing e-waste, showing that server-grade hardware doesn't need to come from a data center.

Technical view

The project converts a retired smartphone into a general-purpose server, likely via a Linux userland such as Termux or UserLAnd, or by flashing a Linux distro like postmarketOS, then exposing standard services over Wi-Fi or LTE. Phones offer ARM SoCs with solid performance-per-watt, an integrated UPS via battery, and optional cellular failover, making them an attractive low-power always-on node. It's replicable by anyone with a spare Android device and either a Termux-based environment or full OS reflashing via bootloader unlock.

Hacker News · 437 ptsBuildable

A physicist rigged his pet hamster’s wheel to upload to Strava

A physicist wired his hamster's exercise wheel to log its nightly runs on Strava.

Strava is an app where athletes track and share runs, rides, and workouts. In this playful project, a physicist attached sensors to his pet hamster's exercise wheel so every spin gets measured, counting rotations and calculating distance from the wheel's circumference, then automatically uploads that data as a workout to Strava, as if the hamster were a tiny athlete. It's a fun, low-stakes hardware hack combining basic physics with electronics tinkering and a real fitness-app integration. It's a charming example of how accessible maker tools have become, letting anyone with basic electronics skills wire up a whimsical passion project and share it with the world.

Technical view

The build likely uses a rotary encoder or a magnet-and-reed-switch sensor on the wheel wired to a microcontroller such as an ESP32 or Arduino, counting revolutions and computing distance from wheel diameter, then pushing the resulting activity to Strava's API to log a workout. It's a straightforward embedded-sensing plus REST API project, replicable with off-the-shelf hall-effect sensors, a Wi-Fi-capable MCU, and Strava's OAuth-based activity-upload endpoint — a good template for learning sensor-to-cloud data pipelines.

Hacker News · 436 ptsConceptual

DeepMind's WeatherNext model achieves breakthrough forecasting cyclones

Google DeepMind's new AI weather model reportedly gets much better at predicting hurricanes.

WeatherNext is an AI system from Google DeepMind designed to forecast weather, and this reports a major leap in predicting cyclones — also called hurricanes or typhoons — including their path, strength, and timing. Traditional forecasting relies on physics-based simulations that are hugely computationally expensive and slow to run; DeepMind's approach instead trains on huge amounts of historical weather data to spot patterns and predict what happens next, often faster and more accurately than classical models. This matters because better cyclone forecasts mean more accurate evacuation warnings, less wasted preparation for storms that miss an area, and potentially saved lives and reduced damage from these destructive storms.

Technical view

WeatherNext is DeepMind's machine-learning weather forecasting system, building on prior work like GraphCast and GenCast, trained on reanalysis and observational data to produce probabilistic forecasts far faster than traditional numerical weather prediction models that solve fluid-dynamics equations on supercomputers. The reported breakthrough is improved skill specifically on tropical cyclone track and intensity forecasting, historically a hard case for both physics-based and ML models due to storms' small spatial scale and rapid intensification. Practitioners could build on this via any published DeepMind models or weights, or through data flowing into operational forecast products, using probabilistic ensemble outputs to improve downstream risk and warning systems.

Hacker News · 429 ptsConceptual

_for-sale DNS records

Expired websites quietly get replaced by spammy 'buy this domain' pages — and nobody notices.

When a company stops using a subdomain or cloud address but forgets to remove the DNS record pointing to it, someone else can grab that abandoned spot and put whatever they want there — often a sleazy 'this domain is for sale' page, or worse. This happens because DNS, the internet's address book, doesn't automatically clean up stale entries, so a link that used to lead to a trusted service can silently start pointing to a stranger's server. It matters because attackers actively hunt for these forgotten records to hijack a trusted company's subdomain for phishing or scams. This piece is essentially a survey of real DNS records caught in that 'for sale' limbo, showing how common and overlooked the problem is.

Technical view

The catalog documents dangling DNS records — CNAMEs or A records left pointing at deprovisioned cloud resources (S3 buckets, CDN endpoints, SaaS subdomains) — that now resolve to generic 'domain for sale' or unclaimed-resource pages. This is the classic subdomain-takeover surface: an attacker can claim the vacated resource and serve content under the victim's trusted domain, enabling phishing, cookie theft, or CSP/CORS bypass. Practitioners can use tools like dnsrecon, subjack, or can-i-take-over-xyz fingerprint lists to audit their own zones for similar orphaned records before an attacker does.

Hacker News · 421 ptsConceptual

Assembly Hall of Shame

A running list of the worst, most embarrassing machine-generated assembly code ever produced.

Compilers translate human-friendly code into raw assembly instructions the processor actually runs, and usually they do a pretty good job of it. This page collects examples where that process goes hilariously wrong — where a compiler produces bloated, redundant, or nonsensical assembly for a task that should be trivial. It's part running joke, part diagnostic tool: by naming and shaming these failures, it helps compiler developers spot patterns worth fixing and helps curious programmers see just how much can go sideways between source code and the chip. It matters because inefficient generated code quietly costs real performance and battery life across billions of devices.

Technical view

The collection catalogs pathological compiler codegen — cases where GCC, Clang/LLVM, MSVC, or similar emit needlessly long instruction sequences, redundant register shuffles, or missed peephole optimizations for simple operations. Each entry typically pairs offending source with generated assembly, often cross-referenced via Compiler Explorer/godbolt, so the regression is directly reproducible. It's useful as a bug-report feeder for compiler backends and as a teaching resource for instruction selection, register allocation, and optimization-pass ordering.

Hacker News · 419 ptsConceptual

Timeline of the OpenAI accidental attack against Hugging Face

OpenAI's own web crawler accidentally flooded Hugging Face's servers like a mini cyberattack.

Big AI companies run automated bots that crawl the web to gather data or check content, and Hugging Face is a hugely popular hub where people share AI models and datasets. This piece walks step by step through how a bot connected to OpenAI ended up sending an overwhelming number of requests to Hugging Face's servers — enough to look and feel like a deliberate denial-of-service attack — even though it was apparently unintentional. It matters because it's a reminder that the automated systems powering AI development can accidentally strain shared infrastructure the whole industry depends on. The timeline format suggests it reconstructs exactly what happened and when, likely from logs and public statements.

Technical view

The writeup reconstructs, chronologically, how automated request traffic attributable to OpenAI infrastructure spiked against Hugging Face endpoints, causing service degradation resembling a distributed denial-of-service event. Such incidents typically trace back to misconfigured crawler concurrency, retry storms without backoff, or scraping jobs lacking rate limiting against a third-party API. The actionable lesson for practitioners is defensive: rate-limit and fingerprint high-volume automated clients, and if you operate crawlers, implement exponential backoff and cap concurrent connections per target host.

Hacker News · 401 ptsConceptual

The Nixpkgs core team has disbanded

The small group steering one of Linux's biggest software repositories just walked away.

Nixpkgs is the huge, community-run collection of software packages that powers the Nix package manager and NixOS operating system — think of it as a giant shared toolbox thousands of projects rely on. A 'core team' is the small group of trusted maintainers who make final calls on tricky decisions and keep the project coherent. This post reports that team has dissolved, raising real questions about who's now in charge of resolving disputes or merging big changes. It matters to anyone using Nix/NixOS because leadership vacuums in major open-source projects can stall development or eventually lead to a fork if the community can't agree on what comes next.

Technical view

Nixpkgs, the package set backing NixOS and the Nix package manager, has apparently lost its formal core/steering team, the governance body that historically handled escalations, RFC approval, and infrastructure decisions for the roughly 100k-package monorepo. This mirrors governance crises in other large volunteer-run open-source projects, where maintainer burnout and unclear authority structures create a vacuum. Downstream implications include potential slowdowns in merging PRs or making breaking changes to nixpkgs-unstable/stable channels — worth monitoring the project's GitHub discussions or RFC repo for whatever structure replaces it.

Hacker News · 400 ptsBuildable

How I use LLMs to learn complex topics

A step-by-step personal method for turning ChatGPT-style AI into a genuine tutor for hard subjects.

Large language models (LLMs) like ChatGPT can explain things, but just asking them questions doesn't automatically make you understand a topic deeply — you need a deliberate approach. This post shares one person's practical technique for using an AI chatbot as a learning partner: things like asking it to quiz you, explain a concept multiple ways, or catch gaps in your understanding. It matters because AI is reshaping self-education, and a thoughtful method can be the difference between shallow 'did I get the right answer' chats and real mastery of a hard topic like math or a new programming language. It's essentially a study-skills guide adapted for the AI era.

Technical view

The post likely details a repeatable workflow for using LLMs as a learning scaffold — techniques such as active-recall prompting (asking the model to quiz rather than lecture), requesting multiple explanatory framings of the same concept, or using the model to generate practice problems of graded difficulty while iteratively drilling into prerequisite gaps it surfaces. A practitioner could replicate this by turning it into reusable prompt templates or a small tool that manages spaced-repetition-style follow-ups against a knowledge tree for a given subject. The value-add over naive Q&A chatting is the explicit pedagogical structure imposed on the interaction.

Hacker News · 371 ptsConceptual

Hardware backdoors in some x86 CPUs

Some Intel-compatible chips reportedly hide secret backdoor instructions built right into the hardware.

Modern CPUs are incredibly complex, with millions of hidden features even most programmers never see, and occasionally researchers discover secret, undocumented functionality baked into the chip itself — a 'hardware backdoor.' This piece reports on x86 processors, the family of chips used in most PCs, that reportedly contain such a hidden mechanism, potentially letting someone bypass normal security protections directly at the hardware level, below the reach of any operating system or antivirus. It matters enormously because software security is meaningless if the chip underneath can be silently subverted — there's no patch that fixes hardware designed this way. Findings like this usually come from researchers reverse-engineering chip firmware or microcode.

Technical view

The report concerns undocumented, privileged functionality embedded in certain x86-compatible CPUs, echoing prior disclosures like the VIA C3 'Rosenbridge' backdoor, where an alternate instruction set could be unlocked to jump straight to ring-0 execution from user mode. Such backdoors are typically found via reverse-engineering of microcode, JTAG/debug interfaces, or leaked documentation, and their danger is operating below the OS and hypervisor, invisible to conventional security tooling. The actionable takeaway for practitioners is supply-chain and hardware-trust auditing — verifying microcode signing, disabling undocumented debug instructions where possible, and treating CPU firmware updates as security-critical.

Hacker News · 358 ptsBuildable

Dithered QR Codes

QR codes that double as grainy, stylish photos you can still scan with your phone.

A QR code is that black-and-white grid your phone camera scans to open a link, and normally it looks purely functional and ugly. Dithering is an old image trick — used since early black-and-white printing — that fakes shades of gray using patterns of tiny black and white dots. This technique blends the two: it takes a real photo, dithers it into a similar pattern of dots, and cleverly arranges them so the result is simultaneously a recognizable picture and a fully scannable QR code. It matters as fun applied creative engineering — turning a purely utilitarian barcode into an artistic design element, like on a poster or album cover, that quietly hides a working link.

Technical view

The technique blends a QR code's required black/white module pattern with a dithered halftone rendering of an arbitrary image, exploiting QR error-correction tolerance (typically level M/Q/H, allowing roughly 15-30% of modules to be 'wrong' and still decode) so dithered pixels can deviate from the strict pattern while remaining scannable. Implementation usually involves generating the QR matrix at a target error-correction level, then applying an error-diffusion dither (e.g., Floyd–Steinberg) constrained per-module toward the correct QR bit, and verifying scannability by re-decoding the output. A practitioner could build this with an image library plus a QR-encoding library, tuning error-correction level and dither threshold to balance visual fidelity against reliable scanning.

Hacker News · 357 ptsRunnable

Windows 11's built-in Weather app wastes more than 1 GB of RAM

Windows 11's simple weather widget is somehow hogging over a gigabyte of memory.

The little Weather app built into Windows 11, the one showing temperature in the taskbar, should be lightweight — it's just displaying a few numbers and icons. Instead, this report finds it consuming more than a gigabyte of your computer's memory, an absurd amount for something so simple, which can slow down the rest of your machine. It matters because it's a symptom of a broader trend: modern operating systems bundle web-technology-based mini-apps for even trivial features, and that convenience for developers costs users real performance. It's the kind of bloat complaint that resonates with anyone who's felt their PC get sluggish for no obvious reason.

Technical view

The report documents the Windows 11 Weather widget/app process consuming upwards of 1GB of RAM, disproportionate to its UI complexity, consistent with it being implemented via a Chromium/WebView2-based wrapper rather than a native lightweight component. This pattern — using web rendering engines for simple system widgets — trades development convenience for significant memory and startup overhead, a known criticism of Windows 11's revamped Widgets board and taskbar weather integration. Practitioners troubleshooting this can check Task Manager for the responsible process (often Widgets.exe or a WebView2 host) and mitigate by disabling the taskbar weather widget via taskbar settings or group policy.

Hacker News · 350 ptsConceptual

U.S. Department of Energy Launches the Genesis Open Models Initiative

The US Department of Energy just launched its own open AI models for anyone to use.

The Department of Energy — the US agency that runs national science labs and supercomputers — has started a program to build and release AI models with their inner workings made public, rather than locked away by a private company. The idea is to give scientists, researchers, and the public access to powerful AI tools built with taxpayer-funded computing power, so progress isn't controlled by a handful of tech firms. This matters because it could make cutting-edge AI more transparent and available for things like scientific discovery, not just commercial products.

Technical view

DOE is standing up an initiative to develop and release open-weight AI models, likely leveraging its national lab supercomputing infrastructure (e.g., systems at Oak Ridge or Argonne) for training. The stated goal is public, inspectable models rather than closed commercial ones, positioning government compute resources as a counterweight to industry-controlled frontier models. Details on architecture, scale, or specific labs involved aren't given, but practitioners should watch for open licensing terms and domain focus (likely scientific/research applications) as it develops.

Hacker News · 328 ptsBuildable

We replaced Redis with MySQL for inventory reservations and it scaled

A team ditched Redis for plain MySQL to reserve inventory — and it scaled better.

When an online store needs to briefly 'hold' an item in your cart so two shoppers can't both buy the last one, engineers often reach for Redis, a fast in-memory database built for speed. This team instead used MySQL, a standard, more old-fashioned database, to manage those temporary holds — and found it handled the load just fine while being simpler to run. The appeal is that MySQL naturally guarantees strong consistency (no accidentally overselling stock) using database transactions, whereas juggling two separate systems adds complexity and potential bugs. It's a reminder that boring, well-understood tools can outperform trendier ones once you actually measure.

Technical view

The team replaced a Redis-based reservation layer with MySQL row-level locking and transactions (e.g., SELECT FOR UPDATE or optimistic concurrency with version columns) to manage inventory holds during checkout. This trades Redis's raw in-memory speed for MySQL's native ACID guarantees, eliminating cross-system consistency bugs (e.g., Redis and the source-of-truth DB drifting) and reducing operational surface area. Practitioners considering this should benchmark lock contention under peak concurrency and ensure proper indexing on the reservation table, since correctness-by-default here trades off some raw throughput headroom.

Hacker News · 324 ptsConceptual

Retraction: The App Store Rejection of the Week That Was a Correct Rejection

A viral 'Apple unfairly rejected my app' story turns out — Apple was right.

Someone previously posted a popular story claiming Apple's App Store review team rejected their app unfairly, and it spread as another example of Apple's review process being arbitrary or heavy-handed. This follow-up post walks that back, admitting the rejection was actually justified according to Apple's rules. It's a small but useful case of a community correcting a viral narrative once more facts came out, and a reminder to be skeptical of one-sided complaint stories before piling on.

Technical view

This is a retraction thread following up on an earlier post that framed an App Store rejection as capricious; the correction concedes the rejection complied with Apple's guidelines. There's no new technical content beyond the community self-correction — it's a case study in verifying vendor-blame narratives against actual policy text before amplifying them.

Hacker News · 304 ptsConceptual

US Military's cyber command unit grapples with cluster of deaths by suicide

A cluster of suicides is hitting an elite unit of the US military's cyber command.

US Cyber Command runs the military's offensive and defensive hacking operations — high-stakes, secretive work. Reporting describes a troubling cluster of suicide deaths within one of its units, prompting concern and scrutiny over how the military supports the mental health of people doing intensely stressful, isolating, classified work. This kind of story matters because it raises questions about oversight, workplace culture, and whether institutions built for secrecy are equipped to catch and help people in crisis.

Technical view

The report (behind an archive link) describes multiple suicide deaths clustered within a specific US Cyber Command unit, prompting internal review. Specifics on causes, unit identity, and remediation steps aren't detailed in the abstract; the underlying story likely examines factors like security-clearance-driven isolation, operational tempo, and gaps in mental health support structures within classified environments.

Hacker News · 290 ptsRunnable

Lost my phone at the office. Claude suggested tracking Bluetooth signal strength

Lost a phone at the office — Claude suggested tracking it down via Bluetooth signal strength.

Someone misplaced their phone at work and asked the AI assistant Claude for help finding it. Claude suggested a clever trick: use another device to scan for the phone's Bluetooth signal and use how strong or weak that signal is to guess how close you are, essentially playing a 'hot and cold' game with radio waves. It's a nice example of AI being useful for everyday physical-world problems, not just writing code or answering trivia.

Technical view

The suggested technique uses Bluetooth Low Energy signal strength (RSSI) scanning from a nearby laptop or device to estimate proximity to the lost phone — walking around while watching signal strength change to triangulate location manually. This is replicable on Linux via tools like `bluetoothctl` or `hcitool lescan`, and mirrors the proximity-based logic behind consumer 'Find My' features, just done manually and guided step-by-step by an LLM.

Hacker News · 256 ptsRunnable

Ancient Library – 1,060 Greek/Latin texts, click any word to parse it

Click any word in an ancient Greek or Latin text and instantly see its grammar.

This is a digital library of over a thousand classical Greek and Latin texts — think Homer, Cicero, that kind of thing — with a neat trick: click on any single word and it breaks down its grammar for you, like what tense a verb is in or what case a noun takes. That's a huge help for students or hobbyists trying to read ancient languages, since those languages pack a lot of meaning into word endings that are hard to memorize. It turns a giant wall of unfamiliar text into something you can explore and actually learn from as you go.

Technical view

The site pairs a corpus of 1,060 Greek and Latin texts with per-token morphological analysis, likely built on lemmatization and tagging approaches similar to Perseus or CLTK (Classical Language Toolkit), surfacing case, tense, mood, and part-of-speech on click. This is a useful reference or base for anyone building computational classics tools, text annotation pipelines, or language-learning software for classical languages.

Hacker News · 243 ptsBuildable

Os8088: A powerful Mac-like OS for the IBM XT, 286, 386

A brand-new, Mac-like operating system built from scratch for 1980s IBM PCs.

This is a hobbyist project building an entirely new operating system — the software that runs everything else on a computer — designed to look and feel like a classic Mac, but made to run on decades-old IBM PC hardware from the 1980s and 90s (the 8088, 286, and 386 chips). It's an impressive feat of squeezing a modern-feeling graphical interface into machines with tiny amounts of memory and processing power by today's standards. Projects like this matter to retrocomputing fans who love keeping old hardware alive and pushing it further than it was originally designed to go.

Technical view

Os8088 is a from-scratch OS targeting x86 real-mode hardware across the 8088, 286, and 386 generations, implementing a graphical, Mac OS-inspired windowing interface under the severe memory and CPU constraints of that era (e.g., 640KB conventional memory limits, no built-in protected-mode assumptions on the 8088). Building or extending it involves low-level x86 assembly/C work — interrupt handling, device drivers, and a custom windowing/graphics stack — making it a solid reference for OS-dev hobbyists interested in vintage hardware bring-up.

Hacker News · 239 ptsConceptual

Silicon Valley misreads science fiction and undermines democracy

Tech leaders are building the future based on sci-fi warnings they mistook for blueprints.

This is an argument that Silicon Valley executives read science fiction stories — many written as warnings about surveillance, AI takeover, or corporate power run amok — and instead of taking them as cautionary tales, treat them as instruction manuals for what to build. The piece argues this misreading isn't harmless; it shapes real decisions about AI, surveillance tools, and platform power in ways that quietly chip away at democratic institutions. It's a call to read more critically and notice when 'inspired by fiction' has drifted into 'making the dystopia real.'

Technical view

This is a cultural/media criticism essay arguing that tech industry decision-makers' embrace of sci-fi tropes (AI-takeover narratives, techno-utopianism) functions as ideology shaping product design and policy stances, with a specific claim that this dynamic erodes democratic norms. It's an argumentative piece rather than empirical research, useful as a reference point for readers examining the relationship between tech industry culture and speculative fiction.

Hacker News · 220 ptsRunnable

Welcoming the Nepalese Government to Have I Been Pwned

A hacked-password checker just added Nepal's government to its official alert list.

Have I Been Pwned is a free website that tells you if your email or password showed up in a data breach that hackers later leaked online. Governments can register their official domains so that if government employee accounts turn up in a leak, the country gets notified directly instead of finding out from a news article. Nepal has now joined this program, meaning its officials will get faster warnings when their systems are compromised. It matters because breached government credentials are a favorite entry point for attackers trying to break into public infrastructure.

Technical view

Have I Been Pwned (HIBP), the breach-notification aggregator run by Troy Hunt, ingests leaked credential dumps and matches them against registered email domains; national CERTs and governments can subscribe to receive alerts when addresses on their gov domains appear in a new breach. Nepal's government has been onboarded to this program, joining a growing list of national partners that get proactive domain-wide breach notifications. Practically, this lets Nepal's CERT correlate breach exposure with its own asset inventory and prioritize credential resets before compromised accounts are used for lateral movement. Anyone running an org can replicate the underlying idea using HIBP's API or a self-hosted breach-checking tool against their own domain list.

Hacker News · 219 ptsBuildable

Tom Stanton's supersonic trebuchet breaks sound barrier with gravity alone

A YouTuber built a giant slingshot that flings a tip past the speed of sound using only falling weight.

A trebuchet is a medieval-style catapult that launches a projectile by swinging a heavy counterweight, converting the energy of a falling mass into speed at the far end of a long arm. YouTuber Tom Stanton scaled this idea up so the very tip of the throwing arm — the part that snaps forward last, like the tip of a cracking whip — moves faster than sound, producing an actual sonic-boom crack. The engineering trick is balancing arm length, counterweight mass, and a whip-like sling release so all the falling energy concentrates into that tiny tip instead of being wasted. It's a fun, visceral demonstration that you don't need explosives or engines to break the sound barrier — just clever leverage and gravity.

Technical view

The build is a mechanical energy-amplification chain: gravitational potential energy from a large counterweight accelerates a long lever arm, and a trailing sling multiplies tip velocity further, similar to whip-cracking dynamics where wave energy concentrates into a shrinking mass at the tip. By tuning counterweight mass, arm-length ratios, and sling release timing, Stanton pushes the projectile-end velocity past roughly 343 m/s (Mach 1 at sea level), evidenced by an audible sonic crack. This is a practical demonstration of mechanical advantage and energy conservation without any combustion or motor input. Replicating it requires careful trebuchet geometry calculations (arm ratio, counterweight drop height) and high-speed audio/camera verification to confirm the sonic boom.

Hacker News · 217 ptsRunnable

Kitesurf: Agent-first browser that runs in V8 isolates

A web browser built for AI agents, running each page in a lightweight sandboxed JavaScript engine.

Most browsers are designed for humans clicking and scrolling, but Kitesurf is designed for AI agents that need to browse the web autonomously — filling forms, reading pages, clicking links — without a human watching. Instead of a full heavyweight browser engine, it runs inside 'V8 isolates,' the same lightweight, fast-starting sandboxed environments that power services like Cloudflare Workers, letting many isolated browsing sessions spin up quickly and cheaply. The core idea is to strip browsing down to what an automated agent actually needs — DOM access and execution — rather than everything a human-facing browser carries. This matters because as AI agents increasingly need to act on the live web, existing browsers are slow, heavy, and hard to run at scale for that purpose.

Technical view

Kitesurf implements browser page execution on top of V8 isolates rather than a full browser engine like Chromium's renderer, trading rendering fidelity for fast cold-starts, low memory overhead, and easy horizontal scaling of concurrent sessions — the same isolate model used by serverless edge platforms. This architecture targets programmatic/agentic use cases (LLM-driven navigation, scraping, form-filling) where deterministic DOM access and JS execution matter more than pixel-perfect rendering or a visible UI. Developers could build on it by scripting agent workflows directly against isolate instances, potentially achieving much higher session density per host than spinning up headless Chromium instances.

Hacker News · 209 ptsRunnable

Triton: DirectX 11 Driver for QEMU

A new driver lets virtual machines run real DirectX 11 graphics inside QEMU.

QEMU is a popular open-source tool for running a virtual computer (a 'VM') inside your real one, but VMs have historically struggled with games or graphics-heavy Windows software because they lack proper GPU driver support. Triton adds a DirectX 11 driver for QEMU, letting a virtual machine's guest operating system talk to modern graphics APIs and get hardware-accelerated rendering instead of a slow software fallback. The trick is building a paravirtualized graphics driver — essentially a translator that passes the guest OS's DirectX calls through to real GPU hardware on the host. This matters for anyone who wants to run Windows games or GPU-accelerated apps inside a sandboxed or virtualized environment without a big performance hit.

Technical view

Triton implements a DirectX 11 guest driver paired with a QEMU-side backend, translating D3D11 API calls from the virtual machine into commands executable on host GPU hardware — a paravirtualized GPU approach similar in spirit to virtio-gpu/VirGL but targeting DX11 rather than OpenGL. This gives Windows guests hardware-accelerated 3D rendering under QEMU instead of relying on slow software rasterizers like WARP. Practitioners can use it to run graphics-dependent Windows applications or games inside QEMU VMs for testing or sandboxing, and could extend the translation layer to cover more of the D3D11 feature set.

Hacker News · 207 ptsConceptual

Can Intel finally beat ARM on performance per Watt?

Can Intel's chips finally match ARM's famous power efficiency?

For years, ARM-based chips (the kind that power phones and Apple's M-series Macs) have been known for doing more computing work per unit of battery power than Intel's traditional x86 chips, which were built more for raw desktop performance than efficiency. This piece asks whether Intel's newest chip designs can close that efficiency gap — get similar or better performance while using similar or less electricity. The comparison usually comes down to chip architecture choices, manufacturing improvements, and how aggressively a company tunes chips for low power versus peak speed. It matters because power efficiency determines battery life in laptops, heat and cooling costs in data centers, and overall running costs.

Technical view

The discussion centers on Intel's evolving core architecture (hybrid P-core/E-core designs, newer process nodes) versus ARM's inherently RISC-based, efficiency-oriented microarchitecture used across mobile and increasingly server/laptop chips (Apple Silicon, Qualcomm Snapdragon X, AWS Graviton). Performance-per-watt comparisons typically weigh instruction-decode complexity (x86's variable-length CISC decode overhead vs ARM's fixed-length RISC decode), process-node parity, and workload-specific tuning rather than raw clock speed. Whether Intel closes the gap likely hinges on specific benchmarks and workloads rather than a blanket verdict, since process shrinks and microarchitectural efficiency work can narrow but not always eliminate the historical gap. Readers wanting to verify claims should compare independent SPEC or performance-per-watt benchmarks across comparable process nodes.

Hacker News · 207 ptsConceptual

Don't use your phone while you poop

Scrolling on the toilet may be quietly wrecking your backside.

This is health advice warning against a very common habit: bringing your phone into the bathroom and scrolling while sitting on the toilet. The concern is that sitting for the short time it takes to actually go is fine, but getting absorbed in a phone stretches that sitting time much longer than needed, and prolonged sitting on the toilet increases pressure on the veins around the rectum. Over time, doctors say, this straining and prolonged pressure can contribute to hemorrhoids (swollen veins that cause pain and bleeding). The fix is simple: keep bathroom visits short and leave the phone outside.

Technical view

The underlying medical claim is that extended sitting on a toilet — encouraged by phone-scrolling distraction — increases venous pressure in the anorectal region, because the seated toilet position offers less support to pelvic floor veins than standing or squatting, promoting venous pooling and straining. Chronically repeated over months or years, this is linked by clinicians to a higher incidence of hemorrhoidal disease. There's no device or protocol to build here; the actionable takeaway is behavioral — limiting toilet time to the physiological task at hand (often cited as under 5-10 minutes) and avoiding distractions that prolong sitting, alongside standard advice on fiber intake and hydration.

Hacker News · 203 ptsConceptual

Everything you do is being recorded

An archived article argues that constant digital surveillance has quietly become the norm.

This piece, preserved via an archive link (suggesting the original may be paywalled or taken down), appears to argue that near-constant recording and tracking of everyday life — through phones, apps, cameras, and online services — has become so pervasive that most people no longer notice it. The abstract gives little detail beyond the title, but the core warning is about how much of what we do, say, and where we go is logged somewhere by some system or company. The argument likely draws on the sheer number of sensors and data-collecting services in modern life: smartphones, smart devices, browsers, and apps all quietly gather data as a byproduct of normal use. It matters because most of us consent to this recording indirectly, without fully grasping its scale or how that data might be used later.

Technical view

With only a title and an archive.is mirror link available, the underlying claims and methodology can't be confirmed — archive.is links often indicate the source sits behind a paywall or was removed, which is itself notable given the surveillance topic. Broadly, this genre of piece typically catalogs passive data-collection vectors (mobile OS telemetry, ad-tracking SDKs, IoT devices, browser fingerprinting, CCTV/ALPR networks) and argues that aggregated metadata is as revealing as content. Readers wanting to verify specifics should treat any statistics or case studies in the original as unconfirmed until read directly, and cross-reference with established privacy research (e.g., EFF, Privacy International) rather than assuming the archived text is exhaustive.

Hacker News · 201 ptsConceptual

Responding to the next frontier of critical cyber capabilities

A look at how defenders must evolve as cyberattack capabilities keep escalating.

This appears to be a piece — likely from a security company or government body — about how organizations need to respond as cyberattack tools and techniques keep getting more sophisticated ('the next frontier'). Without more detail in the title alone, the general theme is that as attackers gain new capabilities, defenders need matching new strategies, policies, or technical measures to keep up. This likely involves some mix of updated detection tooling, information sharing between organizations, and possibly regulatory or coordinated response frameworks. It matters because the gap between attacker and defender capability is a constant arms race, and falling behind has real consequences for critical infrastructure and everyday users.

Technical view

Given only the title, specifics of the proposed response framework, threat model, or named capabilities aren't available, so this should be read as a policy/strategy piece rather than a technical disclosure until the source is reviewed directly. Titles like this typically originate from cybersecurity vendors, CERTs, or government cyber agencies discussing emerging threat categories — plausibly AI-enabled attack tooling, supply-chain compromise, or critical-infrastructure targeting — and propose organizational or coordinated responses such as threat-intelligence sharing, updated detection baselines, or regulatory guidance. Practitioners should treat this as a pointer to read the full source for concrete indicators or recommendations, cross-referencing recent advisories from bodies like CISA or NCSC to contextualize which 'frontier' is being addressed.

Hacker News · 199 ptsConceptual

The original URL for this prediction will no longer be available in 11 years (2011)

A 2011 prediction bet that its own web link would vanish within 11 years.

In 2011, someone made a prediction and wryly noted that the very webpage hosting it probably wouldn't survive 11 years. It's a small, self-aware joke about how fragile the internet actually is — sites get redesigned, companies shut down, and URLs (web addresses) quietly stop working even when the ideas behind them still matter. The 'problem' is link rot: the web's memory is far less permanent than people assume. It's a tiny, personal example of a much bigger issue in how we preserve digital information.

Technical view

This is a self-referential 2011 prediction wagering that its own hosting URL would go dead within an 11-year window, illustrating the well-documented phenomenon of link rot, where a significant fraction of web-cited resources become unreachable within a decade. It's an anecdotal data point rather than a formal study, useful mainly as a discussion trigger about URL persistence and archival practices. Mitigations practitioners typically reach for include the Internet Archive's Wayback Machine and persistent identifier schemes like DOIs.

Hacker News · 199 ptsBuildable

Show HN: textlog – A quiet, text-only microblogging platform, open-source, no JS

A stripped-down microblog you host yourself — plain text, zero JavaScript.

textlog is a small, open-source tool for posting short updates online, like a blog crossed with Twitter, but deliberately minimal — no animations, trackers, or JavaScript running in your browser. It addresses the fact that most social platforms are bloated, ad-driven, and slow, so this strips things back to just text on a page that works even with scripts disabled. Because it's open-source, anyone can download the code, run their own copy, and modify it. It's aimed at people who want a calm, ownership-first alternative to algorithmic social feeds.

Technical view

textlog is a self-hostable, open-source microblogging platform built to render entirely as static or server-rendered HTML with no client-side JavaScript, prioritizing simplicity, speed, and minimal attack surface over rich interactivity. This no-JS constraint limits it to standard form posts and full page reloads rather than SPA-style dynamic updates, but yields fast load times and easy deployment on minimal infrastructure. Developers can fork the repo to add features like feeds or federation while preserving the no-JS design discipline.

Hacker News · 194 ptsRunnable

Open-source interactive map for the Aug 12 total solar eclipse

An open-source map lets you pinpoint exactly where totality hits on Aug 12.

On August 12, a total solar eclipse sweeps across part of the Earth, and this project is a map you can explore in your browser to see exactly where and when the moon's shadow brings full darkness. Instead of a static picture, it's interactive, so you can likely zoom into your own location to check timing and how much of the sun will be covered. It's open-source, meaning the code behind the map is public for others to reuse or improve. That matters for anyone planning to watch the eclipse, since precise path and timing data is the difference between seeing full totality or just a partial dimming.

Technical view

This is an open-source, interactive web map plotting the path of totality and contact timings for the August 12 total solar eclipse, likely built with geospatial libraries rendering eclipse-path polygons computed from astronomical ephemerides. Being open-source, the underlying path-calculation and map-rendering code can be reused or adapted for other eclipse events or geospatial visualization projects. Developers interested in astronomical geodata pipelines can inspect how location-specific timing data is sourced and rendered.

Hacker News · 185 ptsRunnable

Analyzing data from Silicon Valley ventures and founders prosecuted for fraud

Crunching the numbers on Silicon Valley founders who got busted for fraud.

This project digs into real cases of startup founders and venture-backed companies in Silicon Valley who were later charged or convicted of fraud — schemes where companies exaggerated their technology or misled investors. It explores how and why fraud happens in a culture that often rewards bold claims and 'fake it till you make it' attitudes. The approach is data-driven: gathering records of prosecutions and analyzing patterns across them, like company types, amounts of money involved, and how the fraud was uncovered. It matters because spotting these patterns could help investors and regulators catch fraud earlier next time.

Technical view

This project compiles and analyzes a dataset of Silicon Valley startups and founders formally prosecuted for fraud, likely aggregating case details such as charges, funding raised, investors, and outcomes from public legal and press records into a structured dataset. The analytical value lies in identifying recurring structural or behavioral red flags, such as valuation inflation or disclosure gaps, across cases rather than treating each scandal in isolation. Others could extend the dataset with more cases or cross-reference it against funding databases like Crunchbase to build predictive fraud-risk indicators.

Hacker News · 184 ptsConceptual

There Are Magic Hexagons of Every Order

Mathematicians confirm hexagon versions of magic squares exist at every possible size.

A 'magic square' is a grid of numbers where every row, column, and diagonal adds up to the same total — this is about its hexagonal cousin, where numbers fill a hexagon shape and every straight line through it sums equally. Mathematicians have long wondered which sizes, or 'orders,' of these magic hexagons can actually be built. This result shows that a valid magic hexagon can in fact be constructed for every order, not just a rare few. It's a satisfying piece of pure mathematics — the kind of elegant, surprising structural pattern that number enthusiasts love, even without a direct practical use.

Technical view

The piece addresses magic hexagons — hexagonal arrays of consecutive integers where every straight-line path sums to a constant magic number — and presents a general construction proving such an arrangement exists for every order n, extending beyond the famously unique order-3 case long treated as a special curiosity. The core contribution is an explicit constructive method or formula generating valid hexagons at arbitrary order. Readers in combinatorics could replicate the construction algorithm to generate and verify magic hexagons computationally for specific orders.

Hacker News · 182 ptsConceptual

Cool URIs Don't Change (1998)

Tim Berners-Lee's 1998 rule: a web address should never break, ever.

This classic 1998 essay by the inventor of the World Wide Web argues that once you publish a web address, you owe it to everyone who links to it to keep that address working forever, or redirect it properly if things move. It tackles 'link rot': when sites reorganize carelessly, old links break, citations die, and the web's shared memory erodes. The approach isn't software but a set of design principles for structuring URLs so they don't need to change even as a site evolves underneath them. It remains one of the foundational ideas in web design, still cited whenever people debate good URL practices or digital preservation.

Technical view

'Cool URIs Don't Change' is Tim Berners-Lee's foundational web-architecture note arguing URIs should be stable, technology-agnostic identifiers decoupled from server implementation details like file extensions, session IDs, or org-chart-based paths, with server-side redirects handling any necessary restructuring. It lays out concrete practices, such as avoiding embedding file type or authorship info in paths and using HTTP 301 redirects to preserve old links when content moves. It remains a canonical reference in API and REST resource naming and is frequently cited in discussions of link rot and persistent identifiers.

Hacker News · 179 ptsConceptual

Taxi drivers rarely die of Alzheimer's

Cab drivers who navigate for a living seem strangely protected from Alzheimer's.

This is about an observation that people whose jobs demand constant, complex spatial navigation, like taxi drivers, appear to die of Alzheimer's disease less often than people in most other professions. Alzheimer's often first damages the hippocampus, the brain region responsible for memory and spatial navigation. This connects to earlier research showing London taxi drivers' hippocampi physically grow as they memorize the city's street layout, suggesting constant navigational exercise keeps that brain region robust. It hints, though doesn't prove, that heavy sustained use of spatial memory might help protect against this specific kind of brain decline.

Technical view

This finding, drawn from occupation-based mortality data, reports Alzheimer's disease as an unusually rare listed cause of death among taxi and ambulance drivers compared to other occupations, aligning with prior neuroimaging work, notably Maguire et al.'s studies of London taxi drivers, showing hippocampal gray matter enlargement from real-time spatial navigation demands. The mechanistic hypothesis is that sustained engagement of hippocampal-dependent circuits confers resilience against the hippocampal atrophy characteristic of early Alzheimer's pathology. It's observational and correlational, so it can't establish causation, but it supports cognitive-reserve and use-dependent plasticity as areas for controlled follow-up study.

Hacker News · 173 ptsConceptual

Melatonin impairs morning cognition in healthy young adults (2023)

That melatonin gummy might leave your brain foggier the next morning.

Melatonin is a hormone supplement many people take to fall asleep faster, but this 2023 study looked at what it does to thinking ability the following morning in healthy young adults. Researchers gave participants melatonin before sleep, likely compared against a placebo, then tested mental sharpness like attention or reaction time after waking. The finding was that melatonin measurably hurt next-morning cognitive performance compared to not taking it. It matters because melatonin is widely treated as a harmless natural sleep aid, and this suggests a real trade-off between falling asleep faster and feeling groggy the next day.

Technical view

This 2023 study administered melatonin to healthy young adults and assessed cognitive performance the following morning using standardized test batteries, comparing results against a control or placebo condition. The reported outcome is measurable impairment in morning cognitive function following melatonin use, implicating residual pharmacological effects such as prolonged receptor activity or circadian phase-shifting extending beyond the sleep period. This raises dosing-timing and half-life considerations for melatonin use, and suggests follow-up work should isolate which cognitive domains, such as attention or executive function, are most affected and at what doses.

Hacker News · 170 ptsConceptual

Gentoo bugzilla closed due AI bot scraper overload

AI bots hammered Gentoo's bug tracker so hard the maintainers had to shut public access down.

Gentoo is a community-run version of the Linux operating system, and like most open-source projects it keeps a public 'bugzilla' site where anyone can report and discuss software bugs. Recently that site got hit by a flood of automated web crawlers, many belonging to AI companies hoovering up text to train language models, sending so many requests that the servers buckled under the load. Rather than keep fighting an arms race against ever-more-aggressive scrapers, the maintainers closed off open access to protect the service for real users. It's a small but telling sign of a bigger problem: AI training crawlers are increasingly straining the free infrastructure that open-source communities depend on.

Technical view

Gentoo's Bugzilla instance experienced sustained high-volume traffic from AI/LLM training crawlers that ignored or circumvented robots.txt and rate-limiting, degrading service for legitimate contributors. The maintainers' response was to restrict or gate public access rather than continue absorbing the load, joining a growing list of open-source infra (SourceHut, some Git forges, Wikimedia) that have reported similar scraper-driven outages. Practitioners running public bug trackers or wikis should consider proof-of-work challenges (e.g., Anubis), aggressive per-IP/ASN rate limiting, and caching layers as mitigations, since IP blocking alone is increasingly ineffective against distributed scraping fleets.

Hacker News · 166 ptsConceptual

Ask HN: What are you working on? (August 2026)

Hacker News' monthly ritual: thousands of strangers quietly compare notes on their side projects.

This is a recurring monthly thread on Hacker News, a popular tech news and discussion forum, where anyone can post a comment about whatever they're building, learning, or curious about right now — no polish or pitch required. It works like a giant, informal show-and-tell: some comments are one-liners about a weekend hack, others are deep dives into a startup idea or a research rabbit hole someone fell into. The format matters because it lowers the bar for sharing — you don't need a finished product or a big audience, just something you're genuinely into. It matters for the community because it surfaces early-stage ideas and unusual interests long before they'd ever make front-page news.

Technical view

This is HN's recurring 'Ask HN: What are you working on?' thread, a community ritual with no formal structure beyond a monthly cadence and open-ended prompt. There's no dataset or claim to evaluate — the value is entirely in the aggregated, unmoderated cross-section of what a large technical audience is independently pursuing that month, which makes it a decent low-cost signal source for spotting emerging tools, niche problems, and pre-launch projects if you're willing to skim hundreds of comments.

Hacker News · 163 ptsBuildable

Making difficulty curves in games

Why games get harder in waves, not straight lines, and how designers tune that curve.

A 'difficulty curve' is the plan for how challenging a game feels as you progress — too flat and it's boring, too steep and players quit in frustration. This piece is about the craft of designing that curve: deciding when to ramp up enemy strength, introduce new mechanics, or give players a breather level after a hard boss fight. The 'how' usually involves techniques like alternating tension and relief (hard level, then easy one), gradually layering new skills the player has to combine, and adjusting numbers based on playtesting data rather than gut feel alone. It matters because getting this pacing right is often the difference between a game that keeps people hooked and one that gets abandoned an hour in.

Technical view

The article covers practical game-design techniques for pacing challenge over a playthrough: sawtooth curves (steady difficulty rises punctuated by easier 'relief' sections), skill-layering (introducing one new mechanic at a time before combining them under pressure), and tuning via playtest telemetry like death counts and time-to-clear per section. A practitioner building a game could apply this by instrumenting checkpoints to log failure rates and iteratively adjusting enemy stats or level layout until the curve matches the intended emotional arc, rather than hand-tuning numbers speculatively.

Hacker News · 162 ptsRunnable

Microsoft Word for Windows 1.1a, Native X64 Port

Someone rebuilt 1990s Word for Windows to run as a real modern 64-bit program.

Microsoft Word for Windows 1.1a is an early-1990s version of the word processor, originally written for 16-bit Windows and Intel x86 chips of that era. A 'native x64 port' means someone took that decades-old source or binary and reworked it to run directly as a modern 64-bit application, rather than through an emulator or compatibility layer. Doing this is tricky because old software makes assumptions — about memory addressing, data sizes, and system calls — that no longer hold on today's hardware and operating systems, so the porter has to hunt down and fix each incompatibility by hand. It matters to retro-computing enthusiasts and software historians as a way to keep genuinely ancient software runnable and inspectable on current machines without relying on fragile virtual machines.

Technical view

This is a reverse-engineering/porting effort that recompiles or re-targets Word for Windows 1.1a — originally a 16-bit x86 Windows 3.x binary — to run natively as 64-bit code, likely requiring reimplementation or shimming of 16-bit segmented memory assumptions, old Windows API calls, and possibly disassembly-derived source reconstruction since original build sources are unlikely to be available. For practitioners interested in retro software preservation or reverse engineering, this is a concrete reference for the class of problems involved in porting 16-bit real/protected-mode Windows applications to modern 64-bit environments, including handling of near/far pointers and legacy GDI/COM interfaces.

Hacker News · 161 ptsConceptual

Incentives are for losers

A contrarian case that dangling rewards makes people worse at their jobs, not better.

This essay argues against a common belief in business and management: that the best way to get people to perform is to attach explicit rewards — bonuses, prizes, rankings — to specific outcomes. The core argument is that incentive schemes often backfire, because people start optimizing for the measurable target instead of the actual goal, and lose the intrinsic motivation that made them good at the work in the first place. The 'how' here is more argument than experiment — drawing on examples of gamed metrics and hollowed-out effort to make the case. It matters for anyone designing compensation, KPIs, or reward systems, since it's a caution against Goodhart's-law-style failure: when a measure becomes a target, it stops being a good measure.

Technical view

The piece makes a normative argument in the tradition of Deci & Ryan's self-determination theory and Goodhart's Law: extrinsic incentive structures tend to crowd out intrinsic motivation and induce metric-gaming behavior, so 'losers' (in the author's framing) are those who need external incentive design to perform, while high performers are driven by internal standards that incentives can actually undermine. There's no new empirical study here — it's an opinion essay — so its practical use is as a design heuristic: before adding a bonus or leaderboard, ask whether it will redirect effort toward the metric rather than the underlying goal.

Hacker News · 152 ptsBuildable

Message your other Claude Code sessions

A new trick lets one AI coding session ping another AI coding session directly.

Claude Code is Anthropic's coding assistant that runs in a terminal and can work on software projects somewhat autonomously. Normally, if you have several of these sessions running at once — say, one working on a backend fix and another on a frontend feature — they have no way to talk to each other; you're the only line of communication. This feature lets one session send a message directly to another running session, so they can coordinate, hand off context, or report status without you manually copying information between them. It matters because as people run multiple AI agents in parallel on related tasks, giving those agents a way to communicate reduces the coordination burden on the human supervising them.

Technical view

This introduces inter-session messaging for Claude Code, letting one running agent session send a message to another active session by ID or name, preserving that target session's context rather than spawning a fresh one. Practically, this enables patterns like a coordinator session dispatching sub-tasks to worker sessions and collecting results, or long-running background sessions reporting progress to a supervising session, without the user manually relaying text between terminals. It's a building block for more complex multi-agent orchestration workflows run locally, complementing existing subagent/task-spawning mechanisms by adding cross-session, not just cross-subagent, communication.

Hacker News · 126 ptsConceptual

John C. Lilly on solid state intelligence and the elimination of man (1978)

In 1978, a psychedelic neuroscientist warned that machines might quietly retire humanity.

John C. Lilly was an unconventional scientist known for studying dolphin communication, sensory deprivation tanks, and altered states of consciousness, and later in his career he wrote speculatively about the future of computing and intelligence. This piece captures his 1978 thoughts on what he called 'solid state intelligence' — his term for computer-based minds — and his concern that as these machines got smarter, humans might gradually become unnecessary or sidelined, not through violent conquest but through slow obsolescence. His approach was philosophical and intuitive rather than technical, blending his background in brain science with fairly free-form futurism. It's notable today mainly because it's an early, almost eerily prescient articulation of AI-existential-risk worries, decades before they became mainstream.

Technical view

This is a historical primary-source piece: Lilly's 1978 writings framing then-nascent computing hardware as 'solid state intelligence,' arguing that a self-improving substrate-independent intelligence could eventually render biological humans redundant, prefiguring later AI-alignment and existential-risk discourse (e.g., Bostrom, Yudkowsky) by roughly three decades. It has no technical mechanism to build on — it's speculative philosophy from a neuroscientist better known for dolphin research and consciousness experiments — but it's a useful citation for tracing the intellectual lineage of AI-risk concerns outside the computer science mainstream.

Hacker News · 123 ptsConceptual

The Grid That Doubles the Strength of the Ground

A simple mesh laid under roads and foundations can make weak ground hold twice the load.

When you build a road, foundation, or embankment on soft or loose soil, the ground itself can be the weak link, sinking or shifting under the weight above it. A 'geogrid' is a plastic or fiber mesh laid into the soil that interlocks with the surrounding dirt and gravel, spreading out the load and stopping particles from shifting sideways the way they normally would under pressure. The trick is mechanical, not chemical: the grid's pattern of ribs and openings lets soil particles lock into it, turning a loose pile of dirt into something that behaves more like a reinforced slab. This matters in construction because it can let engineers build safely on ground that would otherwise need to be dug out and replaced with expensive imported fill, cutting cost and time.

Technical view

The article explains geogrid soil reinforcement, where a geosynthetic mesh (typically polymer, in uniaxial, biaxial, or triaxial rib patterns) is embedded within a soil layer to provide lateral confinement and interlock with aggregate particles, increasing the effective bearing capacity and reducing rutting/settlement compared to unreinforced fill. Reported performance gains — effectively doubling load-bearing capacity in some applications — come from the mesh restraining lateral soil movement under vertical load, a mechanism now standard in road-base design, embankments over soft soil, and retaining structures. A civil engineer could apply this by specifying grid aperture size and tensile stiffness matched to the aggregate gradation and expected load per relevant geotechnical design standards (e.g., AASHTO or BS 8006).

Hacker News · 121 ptsBuildable

Reviving a four year old reMarkable 2

An old paper-like e-ink tablet gets dusted off to see if it still holds up.

The reMarkable 2 is a tablet with a screen that looks and feels almost like real paper, built for writing and reading rather than watching videos or browsing. After four years sitting in a drawer, the owner tries to bring it back into daily use, which usually means checking whether the battery still holds a charge, updating software that's fallen years behind, and testing whether the pen still writes smoothly on the screen. It's a small experiment in how well a simple, single-purpose gadget ages compared to phones and tablets that get replaced every couple of years. The story matters as a reminder that not all tech needs constant upgrading to stay useful.

Technical view

The piece documents reviving a reMarkable 2 e-ink tablet after roughly four years of disuse, covering firmware updates, Li-ion battery health checks, and stylus/digitizer calibration. Because the device runs a hackable Linux-based OS accessible via SSH, this kind of write-up typically doubles as a guide to sideloading community software and avoiding update-related bricking. It's a useful reference for anyone assessing the long-term maintainability and repairability of dedicated e-paper hardware versus general-purpose tablets.

Hacker News · 117 ptsConceptual

Gateway 2000's hilariously bad ads in the 90s (Part II)

A nostalgic dig through Gateway 2000's gloriously corny cow-themed 90s computer ads.

Gateway 2000 was a 1990s computer company famous for shipping PCs in boxes printed with black-and-white cow spots, leaning into a quirky farm-and-cowboy theme to stand out from boring beige-box rivals. This piece, a second installment, rounds up more of those dated, over-the-top ads to show how strange and unpolished early personal computer marketing really was. Back then, no company had really figured out how to make computers look cool, so ads leaned on gimmicks and humor instead. It's a fun little time capsule of how a scrappy tech industry sold itself before slick branding took over.

Technical view

A retrospective compilation of Gateway 2000's print and TV advertising archive from the 1990s, continuing an earlier installment, likely drawn from period magazines and archived commercials. It serves as a primary-source-adjacent artifact for anyone researching PC industry branding history, documenting how the mail-order vendor used its cow-spot identity and rural-Iowa positioning to differentiate in a rapidly commoditizing market. Useful raw material for a broader history of consumer computing marketing or brand-evolution case studies.

Hacker News · 110 ptsConceptual

Should you stop cracking your knuckles?

That satisfying knuckle-pop habit finally gets checked against the actual science.

Cracking your knuckles makes that pop sound because of gas bubbles forming and then collapsing in the slippery fluid that cushions your finger joints — not from bone grinding on bone, as many assume. For decades people worried it causes arthritis, so researchers have tried to settle the question by comparing joint health between habitual knuckle-crackers and non-crackers, sometimes even imaging joints in real time as they pop. This piece looks at what that evidence actually shows: whether the habit is harmless, mildly annoying, or genuinely risky. It's worth knowing because it's a habit millions of people do daily without ever confirming if it's safe.

Technical view

The article surveys the physiology of knuckle cracking — cavitation within the metacarpophalangeal joint's synovial fluid, where a sudden pressure drop forms a gas bubble that collapses to produce the audible pop — alongside studies (case-control comparisons, real-time MRI/ultrasound) examining links to osteoarthritis, grip strength, and joint swelling. The body of evidence surveyed generally finds no significant association between habitual cracking and degenerative joint disease, though study quality and sample sizes vary. Useful for readers wanting a citation-grounded answer to a persistent health myth rather than anecdote.

Hacker News · 105 ptsConceptual

FCC moves to ban Lidar-equipped foreign drones from US

US regulators move to keep foreign-made, laser-mapping drones out of American skies.

Lidar is a sensing technology that fires laser pulses to build precise 3D maps, and it's widely used in drones for surveying land, inspecting infrastructure, and helping vehicles navigate autonomously. The FCC, the US agency that regulates communications equipment, is moving to restrict foreign-made drones carrying lidar sensors, echoing earlier national-security crackdowns on foreign drone cameras and telecom gear. The concern is that such hardware, often made by companies with ties to rival governments, could quietly collect sensitive data about US infrastructure or send information back overseas. It matters because it's part of a broader push by the US to control which foreign-made hardware gets embedded in critical mapping and infrastructure systems.

Technical view

The FCC is reportedly extending its equipment-authorization restrictions (building on the framework used against companies like DJI and Huawei under national-security authority) to cover lidar-equipped drones from specified foreign manufacturers. This would block import, sale, and marketing authorization for the affected hardware in the US, pushing survey, mapping, and infrastructure-inspection firms that rely on foreign lidar-drone platforms toward domestic or allied alternatives. Practitioners in commercial drone/lidar fields should watch for the named manufacturers and compliance timeline, since it will directly reshape procurement and the lidar-drone supply chain.