TGViewer
Channel Public Channel
Artificial Intelligence AI News

Artificial Intelligence AI News

@machinelearningresearchnews

We are a community of machine learning enthusiasts/researchers/journalists/writers who share interesting news and articles about the applications of AI.

You will never miss any updates on ML/AI/CV/NLP fields because we post them daily. JOIN NOW
Subscribers
3.49K
Photos
128
Videos
115
Links
1.5K
Recent Posts 20 shown
Post #1562 141
Microsoft just released Microsoft-Decision-1, a decision-scoring model that returns calibrated probabilities instead of text. It is post-trained from Qwen3.5-9B and built for routing, classification, verification and agent control.

The core idea is single-pass decision scoring. You give it a situation, a question and a fixed set of options. It returns one calibrated probability per option, so your app can act, defer or escalate in a single call.

In Microsoft’s 36-benchmark comparison (147,137 questions), it scored 83.5% average accuracy, ahead of Quyet-1.0-Large at 81.9%. It ran at 85 ms p50 latency, 35x quicker than GPT-6 Sol. The trade-off: it is text-only, gives no explanations, and its weights are closed. Calibration (92.2) trails Quyet’s 93.1.

At $0.042 per 1M input tokens with free output, it is priced for agent loops that make thousands of small decisions......

Full analysis: https://marktechpost.com/2026/10/09/microsoft-ai-releases-microsoft-decision-1-a-qwen3-5-9b-decision-scoring-model/

Model: https://ai.azure.com/catalog/models/Microsoft-Decision-1

Announcement: https://commandline.microsoft.com/microsoft-decision-1-model-foundry/
  • πŸ’© 2
Post #1561 410
Meta just open-sourced Rebalancer, the assignment solver that has run resource allocation across Meta for 9+ years. It solves ~40M assignment problems per day across 30+ problem formulations, from shard placement to global traffic routing.

The core innovation is separating the problem spec from the solver. You describe objects, bins, constraints and goals once. Rebalancer compiles that into an expression graph, then solves it with either a parallel local search or a MIP solver (Gurobi, FICO Xpress or HiGHS).

The trade-off is explicit. MIP gives optimal answers but the model grows as O(objects Γ— bins), so Meta's largest problems are too big for any MIP solver. Local search works directly on the graph with an O(objects + bins) neighborhood and evaluates millions of moves per second. Meta uses local search for almost all large problems.

Production numbers: P99 solve time of 12s on 265k objects and 3.2k bins. Problems above 1M objects and 5k bins average 171s, across 3.4k+ runs.

Full analysis: https://www.marktechpost.com/2026/10/06/meta-ai-open-sources-rebalancer-a-c-assignment-solver-that-runs-about-40-million-placement-problems-a-day/

Repo: https://github.com/facebook/rebalancer

Paper: https://www.usenix.org/conference/osdi24/presentation/kumar

Technical details: https://engineering.fb.com/2026/09/21/open-source/rebalancer-generic-high-performance-library-assignment-problems/
  • ❀ 1
Post #1560 447
Mistral AI Releases Mistral Large 4 (Le Chonk): A 1.05T Parameter Multimodal MoE Model

Mistral Large 4, nicknamed Le Chonk, went into public preview today. 1.05T total parameters, 49B active per token, 1M context, native image input. Trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral's own European datacenters.

Here is what stood out to me after reading the release and the docs:


β†’ Only about 4.7% of the weights fire per token
β†’ That is how a 1T class model serves at $1.36 per 1M input and $4.18 per 1M output tokens
β†’ The full 1.05T still has to sit in memory, so plan hardware around total, not active
β†’ 93% on Cybench, 82% on CyberGym-E2E, top 5 on the Artificial Analysis Cyber Index
β†’ 61.7% DeepSWE v1.1, 28.3% Terminal-Bench 4.0, 49.8% on the Coding Agent Index
β†’ API is live today as mistral-large-4

I broke down the architecture, benchmarks and a full comparison against DeepSeek V4 Pro, Kimi K3 and GLM-5.3, with an interactive explainer.

Here is the full analysis: https://www.marktechpost.com/2026/10/06/mistral-ai-releases-mistral-large-4-le-chonk-a-1-05t-parameter-open-weight-multimodal-moe/

Technical details: https://mistral.ai/news/mistral-large-4/
  • πŸ‘ 4
Post #1559 464
[AI Model Family Series #1] We just published the complete story of Alibaba's Qwen: every major model from 2023 to 2026, with dates, key features and sources.

7B open weights in 2023 β†’ 2.4T open weights in 2026. Here's the lineup the story covers

2023
β†’ Tongyi Qianwen (Apr): enterprise beta
β†’ Qwen-7B (Aug): first open weights
β†’ Qwen-VL (Aug): first vision model
β†’ Qwen-72B + Qwen-1.8B (Dec)

2024
β†’ Qwen1.5 (Feb): 0.5B–110B, 32K context
β†’ Qwen2 (Jun): 57B-A14B MoE, Apache 2.0
β†’ Qwen2-Math, Qwen2-Audio, Qwen2-VL (Aug)
β†’ Qwen2.5 (Sep): 18T tokens, 100+ models
β†’ Qwen2.5-Coder + QwQ-32B-Preview (Nov)
β†’ QVQ-72B-Preview (Dec)

2025
β†’ Qwen2.5-VL + Qwen2.5-Max (Jan)
β†’ QwQ-32B (Mar): Qwen claims R1-level reasoning
β†’ Qwen2.5-Omni-7B (Mar)
β†’ Qwen3 (Apr): hybrid thinking, 119 languages
β†’ Qwen3-2507 + Qwen3-Coder-480B (Jul)
β†’ Qwen-Image + Qwen-Image-Edit (Aug)
β†’ Qwen3-Max (Sep): first 1T+ Qwen
β†’ Qwen3-Next-80B-A3B (Sep)
β†’ Qwen3-Omni + Qwen3-VL (Sep)

2026
β†’ Qwen3-Max-Thinking (Jan)
β†’ Qwen3-Coder-Next + Qwen-Image-2.0 (Feb)
β†’ Qwen3.5-397B-A17B (Feb): native multimodal agents
β†’ Qwen3.6-Plus, 35B-A3B, Max-Preview, 27B (Apr)
β†’ Qwen3.7-Max (May) + Qwen3.7-Plus (Jun)
β†’ Qwen3.8-2.4T-A95B (Aug): largest open Qwen
β†’ Qwen3.8-27B + Qwen3.8-Flash (Aug)
β†’ Qwen-Image-2.1 (Sep)

The story also covers what the list can't show: the DeepSeek moment, the 2026 leadership exit, and Qwen's shift from all-open to a two-tier license strategy.

Next: Qwen 4 is in training. Qwen 4.5 and Qwen 5 are projected at 5–10T parameters.

Read the full report: https://www.marktechpost.com/2026/10/04/the-story-of-qwen-alibabas-ai-models-from-7b-to-2-4t/

Which model family should we map next: DeepSeek, Llama, Gemma, Mistral or Kimi? Drop it in the replies.......
  • ❀ 4
  • πŸ‘ 1
  • πŸ€” 1
Post #1558 668
Cloudflare Releases Clef: Open-Weight Decision Models That Return Typed Probabilities Instead of Text

It answers typed questions about an input (yes/no, pick one, rate on a scale) in a single forward pass.

No tokens are generated, so there is no output to parse before your agent acts.

βœ… 2 sizes: Clef (27B) and Clef-flash (9B), open weights on Hugging Face
βœ… Apache 2.0 license, commercial use allowed
βœ… Post-trained from Qwen3.8-27B and Qwen3.5-9B, backbones kept frozen
βœ… 38.8 ms median latency for Clef-flash vs 524.1 ms for Jev
βœ… 94.20 macro-F1 on BANKING77 vs 79.74 for Jev
βœ… Reads text, JSON, images and video, 64K-token context
βœ… Up to 64 questions and 4 images per request on Workers AI
βœ… $0.24 (Clef) and $0.09 (Clef-flash) per million input tokens
βœ… Drop-in compatible with the Jev / System One API

Full analysis: https://www.marktechpost.com/2026/10/01/cloudflare-releases-clef-and-clef-flash/

Model: https://huggingface.co/Cloudflare/clef

Docs: https://developers.cloudflare.com/workers-ai/models/clef/
  • ❀ 5
  • πŸ”₯ 1
Post #1557 684
Google DeepMind just announced Gemini 4 Argon, raising the output limit from 64K to 1M tokens in a single response. It is the first model of the Gemini 4 generation, built for long-horizon coding, enterprise knowledge work, and cyber defense.

The core shift is output depth. With room to generate hundreds of thousands of tokens in one trajectory, Argon can tackle large refactors and multi-step problems without splitting work across turns.

In Google's own comparison, Argon leads GPT-6 Astra and Claude Opus 5.5 on DeepSWE v1.1 (77.9%), the Vals Index (68.9%), and AutomationBench (51.3%). It is not a sweep, though: it trails on FrontierSWE v2, Terminal-Bench 4.0, and OSWorld-2.0.

Access is gated. Argon is rolling out first to cyber defenders in the Fairwind Program, with paid API customers and Google AI Ultra subscribers next. Introductory pricing is $2 input / $10 output per 1M tokens......

Full analysis: https://www.marktechpost.com/2026/09/30/google-deepmind-unveils-gemini-4-argon-with-1m-output-tokens-for-coding-knowledge-work-and-cyber-defense/

Announcement: https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/

Fairwind Program: https://deepmind.google/fairwind-program/
  • πŸ‘ 2
Post #1556 633
Search is now the bottleneck for AI agents, and Perplexity just rebuilt theirs from the ground up.

Photon is their new retrieval and ranking engine, written in Rust. It replaces the open-source engine they had forked for years.

The old engine's problem: the index outgrew RAM. Cold reads caused page faults, p99 sat near 800 ms, and index merges pushed it to ~1.2 s.

Photon's fix, in 4 moves:

1/ Compact posting lists that pick a format per block: inline, sparse arrays, or bitmaps
2/ A budgeted, WAND-like traversal that skips candidates before reading exact scores
3/ Disk reads batched through io_uring, so waits overlap instead of piling up
4/ Index builds moved off the serving nodes entirely

The result in production: p99 dropped to ~65 ms, on ~20% fewer machines, while storing ~2.5x more data per document.

For developers, Photon powers a new Fast Search mode:

β†’ 160 ms p50 / 230 ms p95 per call
β†’ $1 per 1K requests (vs $5 standard)
β†’ ~68% lower estimated agent task cost at comparable quality

Technical details + full breakdown in the reply πŸ‘‡
Full breakdown: https://marktechpost.com/2026/09/30/perplexity-introduces-photon-a-rust-based-retrieval-engine-that-cuts-p99-latency-from-800-ms-to-65-ms/

Technical details: https://perplexity.ai/hub/blog/photon
  • πŸ‘ 3
  • 🀣 1
Post #1555 927
NVIDIA Releases Nemotron 3 Diarization: A 100M-Parameter Open-Weight Model That Tracks 8 Speakers in Real Time

It answers "who spoke when" in a conversation, even when people talk over each other.
One open-weight checkpoint handles both offline audio and real-time streaming.

βœ… 100M parameters, open weights on Hugging Face

βœ… OpenMDW-1.1 license, commercial use allowed

βœ… Tracks up to 8 speakers, overlap included

βœ… 14.72% DER on Diarization-Bench vs 19.3% runner-up

βœ… 4 streaming presets, from 30.4 s down to 0.32 s

βœ… 41.0% average relative DER cut vs Streaming Sortformer

βœ… 15,113x batched RTFx on RTX PRO 5000, 30.4 s preset

βœ… Runs through NVIDIA NeMo on Linux GPUs

Full breakdown + interactive explainer: https://www.marktechpost.com/2026/09/23/nvidia-releases-nemotron-3-diarization/

Model: https://huggingface.co/nvidia/Nemotron-3-Diarization
NVIDIA blog: https://huggingface.co/blog/nvidia/nemotron-diarization
Live demo: https://huggingface.co/spaces/nvidia/nemotron-diarization
  • ❀ 4
  • πŸ‘ 1
Post #1554 813
Nokia AI Research Open-Sources AnyJev: A Training-Free Layer That Turns Any Open LLM Into a Calibrated Decision Model

It reads typed answers straight from the model's next-token distribution. It removes position and label bias without any training.

β–ͺ️ Apache-2.0, pip install "anyjev[hf]"

β–ͺ️ Choice, yes/no and score questions, nothing generated

β–ͺ️ Hugging Face transformers and vLLM backends

β–ͺ️ L0 needs zero labels, L1 needs 100 to 500

β–ͺ️ Order-flip rate 0.230 to 0.073 at L0

β–ͺ️ Calibration error 0.240 to 0.095 at L1

β–ͺ️ Auto-decidable at 5% error: 7.7% to 52.0% at L1

β–ͺ️ Tested on Qwen3-8B, BANKING77 20-way, 300 test items


Accuracy moves 6 points, but the traffic you can safely automate grows 6.8x.

Full analysis: https://www.marktechpost.com/2026/09/23/nokia-open-sources-anyjev-a-training-free-layer-that-turns-any-open-llm-into-a-calibrated-decision-model/

Repo: https://github.com/nokia-applied-research/AnyJev
  • πŸ‘ 1
Post #1553 949
TypeSafe AI released Jev, a model from an RLHF (reinforcement learning from human feedback) co-inventor that does not generate text.

You send it state and typed questions. It returns decisions with probabilities
your code can branch on, so there is nothing to parse.

- Founder Diogo Almeida co-invented RLHF at OpenAI
- $40M seed round led by DCVC
- 3 primitives: Choice, Score, Noul
- Every question in a request scored in parallel
- $0.042 per 1M input tokens, output tokens free
- 70ms to 500ms end to end, vendor reported
- 0.114s vs 8.566s in TypeSafe's own recorded demo
- Vercel: up to 18x faster than GPT Luna, p95
- Listed on Vercel AI Gateway as typesafe-ai/jev

Confidence is the product. Your code sets the threshold where automation stops
and a human takes over.

The 193.6x faster and 444.6x cheaper claims come from TypeSafe's own evals, not independent tests.

Read the full analysis: https://www.marktechpost.com/2026/09/19/typesafe-ai-releases-jev/

Technical details: https://typesafe.ai/blog/introducing-system-one-models-and-jev

Join the waitlist: https://typesafe.ai/
  • ❀ 3
  • πŸ‘ 1
Post #1552 793
Linkup just open-sourced SPARSEUP, a 149M sparse embedding model scoring 56.4 on BEIR-13.

It is a SPLADE-style retriever that outputs readable token weights, not an opaque dense vector. It fills the sparse slot next to LightOn's DenseOn and LateOn, trained on the same open data.


βœ… Apache 2.0, weights on Hugging Face

βœ… ModernBERT backbone, English only

βœ… 56.4 nDCG@10 on BEIR-13, excluding MS MARCO

βœ… Beats splade-v3 (51.7) and opensearch doc-v3-gte (54.6)

βœ… ~380Β΅s per query at >97% recall, Seismic, MS MARCO

βœ… 47 non-zeros per query, 190 per document

βœ… Logit shift of 15 keeps vectors sparse from the start

βœ… Top-12 expansion cap per token, not per vector

βœ… Case folding cuts output from ~50k to ~34k dims


Every dimension is a token, so you can read exactly why a document matched.

With identical backbone and data, it still trails DenseOn by 1.52 points and LateOn by 2.5. Linkup also documents weak expansions, like number queries spilling into unrelated terms.

Full breakdown: https://www.marktechpost.com/2026/09/19/linkup-research-releases-sparseup/

Model: https://huggingface.co/Linkup-Platform/linkup-sparseup-embed-v1

Technical blog: https://www.linkup.so/blog/introducing-sparseup-by-linkup
  • ❀ 2
Post #1551 738
Stanford Researchers Release Paper2Agent: Turning Research Papers Into AI Agents That Reproduce Results and Run on New Data


Paper2Agent turns a research paper and its codebase into a tested AI agent. You don't have to clone repos, install dependencies or debug environments before using a paper's method.

It builds an MCP server from the paper's code, then generates and runs tests against the paper's own reference outputs. Tools that keep failing are excluded, so every tool in the final server has passed validation.

You can connect the server to Claude Code or any MCP-compatible agent and ask it to apply the paper's method to your own data in plain language.

On the AlphaGenome paper, Paper2Agent built 22 validated tools in about 45 minutes for US $14. The resulting agent scored 100% on 15 novel queries, compared with 78.7% for Claude Code with direct repo access and 56.0% for Biomni.

Across 100 computational biology papers from bioRxiv, 74 were agentified and 593 of 599 proposed tools passed validation. On 300 benchmark questions, it reached 91.2% accuracy at US $0.20 per query.

Learn more with the following resources:

πŸ“° Full breakdown: https://www.marktechpost.com/2026/09/16/stanford-researchers-release-paper2agent-turning-research-papers-into-ai-agents-that-reproduce-results-and-run-on-new-data/

πŸ“„ Paper:
https://www.nature.com/articles/s41586-026-11044-y

⭐ GitHub:
https://github.com/jmiao24/Paper2Agent

πŸ€— AlphaGenome MCP server:
https://huggingface.co/spaces/Paper2Agent/alphagenome_mcp

πŸ§ͺ Try it:
https://paper2agent.ai/live
  • πŸ”₯ 2
Post #1550 754
Google released Gemini 3.8 Live and 3.8 Live Extended Thinking, native speech to speech models built for production voice agents.

Both run in the Gemini Live API and replace cascaded ASR plus LLM plus TTS pipelines with a single model. The defining capability: the conversation never pauses while reasoning and tool calls finish in the background.


βœ… #1 on Artificial Analysis Speech to Speech Index, 82.6

βœ… Executes tools and API calls in the background mid conversation

βœ… Auto switches between 97 supported languages mid conversation

βœ… $0.005/min audio input, $0.018/min audio output


Extended Thinking still scores just 35.1% on Sierra's Ο„-Voice-banking benchmark, so hard multi-step voice workflows are far from solved. Weights are not open; access is API only, with enterprise access in private preview.


Full analysis: https://www.marktechpost.com/2026/09/15/google-releases-gemini-3-8-live-and-3-8-live-extended-thinking-for-production-grade-voice-agents/

Technical details: https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-live-gemini-3-8-live-extended-thinking/
  • πŸ‘ 1
  • πŸ”₯ 1
Post #1549 814
NVIDIA Open-Sources OSMO: One YAML Orchestrates Physical AI Training, Simulation, and Robot Testing

Here's how it works. πŸ‘‡

1. One YAML, three computers

Robot pipelines span 3 compute tiers: training on GB200/H100, simulation on RTX PRO 6000, hardware-in-the-loop on Jetson AGX Thor. OSMO treats all 3 as backends of one control plane.

β†’ Tasks name a platform (gb200, rtx-pro-6000, jetson-agx-thor), never a cluster

2. Dependencies are data

The README example chains 3 tasks: Isaac Sim β†’ PyTorch training with 8 GPUs β†’ ROS eval on Jetson, writing results to a named dataset.

β†’ Placement from platform, ordering from inputs, persistence from outputs

3. Same file, laptop to cloud

Runs on KIND locally and on EKS, AKS, GKE, on-prem, or air-gapped clusters. Release 6.3.0 added a multi-provider deploy script with MinIO, Azure Blob, or S3 storage.

β†’ Zero code changes between environments


4. Built for production

NVIDIA KAI Scheduler by default, NVLink topology-aware placement, per-group timeouts, RBAC sidecar, OAuth2 login, TLS at the gateway, cloud workload identity.

β†’ Powers Project GR00T, Isaac Lab, Isaac Sim, and Isaac ROS internally

Full analysis: https://www.marktechpost.com/2026/09/14/nvidia-open-sources-osmo-one-yaml-orchestrates-physical-ai-training-simulation-and-robot-testing/

Repo: https://github.com/NVIDIA/OSMO
  • πŸ”₯ 1
Post #1548 758
A Princeton Researcher Proposes Recurrent Looped Transformer (RLT): a decoder that carries its full state across every prompt and response token, with 96 logical blocks per token and no reset at the boundary.

Here's how it works. πŸ‘‡

1. One state, never reset
A causal encoder builds KV memory in parallel. A recurrent decoder merges each token's encoder feature with the previous final hidden state, then runs sliding-window attention, cross-attention to encoder memory, and an FFN. The state H_t = (s_t, C^D_t) crosses the serving split untouched.
β†’ Proposition 3.1: moving the prompt-response split does not change the conditional distribution

2. Fixed work, growing depth
Reference config ties 48 encoder and 48 decoder layers with shared attention and FFN weights.
β†’ 96 logical blocks per token, fixed
β†’ 48t decoder blocks on the state path after t tokens

3. RL replay under current parameters
The sampler records behavior log-probs. The trainer rebuilds encoder memory, recurrent output, and every SWA cache from H_0 before scoring each action.
β†’ Old rollout states are never reused; weight updates invalidate all caches

The key takeaway: a design spec that makes pretraining, SFT, sampling, and RL replay share one transition.

Full analysis: https://www.marktechpost.com/2026/09/13/a-princeton-researcher-proposes-recurrent-looped-transformer-rlt/

Paper: https://github.com/yifanzhang-pro/recurrent-looped-tranformer/blob/master/Recurrent_Looped_Transformer.pdf

Git Repo: https://github.com/yifanzhang-pro/recurrent-looped-tranformer

Amazing explanation in author's research blog: https://yifanzhang-pro.github.io/recurrent-looped-tranformer/
  • πŸ‘ 1
Post #1547 868
Sparse attention already cut long-context compute. The KV cache sitting in HBM and on SSD is now the bottleneck DeepSeek AI went after, and they cut theirs to 890 bytes per token.

They released DeepSeek-V4.1-Flash, a 552B MoE model with 1M-token context that activates only 8B parameters per token during prefill and 16B during decode. The global KV cache footprint is roughly 1/4 of DeepSeek-V4-Flash and about 437x smaller than DeepSeek-V1, and the weights are open under an MIT license.

Here's what's actually interesting:


β†’ Causal Encoder-Decoder: prompt tokens stop at layer 20, decoder global KV is projected from the encoder's final hidden state, prefill compute nearly halves

β†’ Compressed Sparse Attention 2: three statically assigned layer modes, Reindex reuses main KV and indexer K, Reuse also reuses Top-K indices

β†’ FP4 main KV cache via quantization-aware training, SWA Bounded Replay replays only 128 tokens instead of layers times window

β†’ Terminal-Bench 2.1: 90.6 vs 89.1 for Opus-5

β†’ DeepSWE v1.1: 74.2 vs 74.0 for Opus-5 and 73.0 for GPT-5.6 Sol

β†’ Terminal-Bench 4.0: 31.2 vs 51.8 for Opus-5, so the gap on expert-level science tasks is still real


Full analysis: https://www.marktechpost.com/2026/09/10/deepseek-ai-released-deepseek-v4-1-flash-with-1m-context-fp4-kv-cache-and-cross-layer-attention-reuse/

Technical Details: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash

Paper: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/main/DeepSeek_V41_Tech_Report.pdf
  • 🀯 2
  • πŸ”₯ 1
Post #1546 821
NVIDIA Announces CUDA Rust with cuda-oxide (SIMT) and cutile-rs (Tile) for Compile-Time-Safe GPU Kernels

NVIDIA has announced CUDA Rust, a push to make Rust a first-class language for writing GPU kernels. Rust code could already launch CUDA kernels, but the kernel body usually had to be written elsewhere. CUDA Rust closes that gap with two NVlabs open-source projects: cuda-oxide for the SIMT model and cutile-rs for the newer Tile model. Both compile Rust kernels natively and use Rust’s ownership rules to reject aliasing bugs at compile time.

Here's how it works. πŸ‘‡

1. cuda-oxide (SIMT track)
A custom rustc codegen backend. #[kernel] functions go from Rust MIR through the Pliron IR framework and LLVM down to PTX. Host and device code live in one file.
β†’ Needs a pinned nightly (2026-04-03), CUDA 12.x+, compute capability 8.0+
β†’ Status: early alpha

2. cutile-rs (Tile track)
Each tile block runs the kernel body once as a single logical thread. The compiler owns thread mapping and memory layout. The kernel AST is embedded in the host binary and JIT-compiled through CUDA Tile IR on first launch.
β†’ Stable Rust 1.89+, CUDA 13.3, no nightly, no custom LLVM
β†’ Published on crates.io, already used in Hugging Face's Grout...

3. What the compiler catches
Pass a kernel's output buffer as one of its own inputs and it does not build.
β†’ SIMT: error[E0502]: cannot borrow c_dev as mutable because it is also borrowed as immutable
β†’ Tile: error[E0382]: use of moved value: z
cutile-rs carries ownership across the launch boundary, which NVIDIA calls the stronger guarantee....

Full analysis: https://www.marktechpost.com/2026/09/08/nvidia-announces-cuda-rust-with-cuda-oxide-simt-and-cutile-rs-tile-for-compile-time-safe-gpu-kernels/

Technical details: https://developer.nvidia.com/blog/introducing-cuda-rust-two-tracks-for-writing-gpu-kernels
  • πŸ‘ 2
  • ❀ 1
Post #1545 676
Google DeepMind Releases AlphaGenome Atlas: Precomputed Molecular Effect Predictions and AVI Scores for All 9 Billion Single-Letter DNA Changes in the Human Genome.

No lab assay. No per-variant model run. No coding required to query it.

Here is how it works:

1. Precompute instead of predict on demand AlphaGenome was run across every possible single-nucleotide variant in the human genome and the outputs were stored. Researchers now look up a variant instead of running a model on it.
β†’ 9 billion variants scored
β†’ 1 petabyte dataset, over 30x the size of the AlphaFold Database

2. One score for coding and non-coding DNA The AlphaGenome Variant Impact (AVI) score folds AlphaGenome's regulatory predictions together with AlphaMissense's protein-impact predictions into a single rankable number.
β†’ Works across the 2% coding genome and the 98% non-coding genome
β†’ DeepMind reports best-in-class results on variant pathogenicity and rare disease benchmarks

3. The score is decomposed, not opaque Each AVI score splits into additive feature attributions across categories like chromatin accessibility, splicing, and conservation, so you see which process a variant is predicted to disrupt.
β†’ Paired with a compendium of 2,500+ recurrent DNA motifs and their locations

4. Rare disease result At the Broad Institute, the AVI score reprioritized variants that earlier analyses had missed and surfaced one in DNM1, a gene linked to epileptic encephalopathy. The prediction: an incorrect splice site extending the protein.
β†’ Experimental screens validated it and found nearby variants with similar effects....

Full analysis: https://marktechpost.com/2026/09/08/google-deepmind-releases-alphagenome-atlas-with-precomputed-molecular-effect-predictions-and-avi-scores-for-9-billion-human-dna-variants/

Paper: https://storage.googleapis.com/deepmind-media/DeepMind.com/Blog/alphagenome-atlas-a-predictive-map-of-every-possible-dna-letter-change-in-the-human-genome/alphagenome-atlas.pdf

Technical details: https://deepmind.google/blog/alphagenome-atlas-a-predictive-map-of-every-possible-dna-letter-change-in-the-human-genome/
  • ❀ 1
Post #1544 734
H Company Releases NeoMME: A Family of 260M and 800M Single-Tower Multimodal Encoders That Drop the Vision Tower and Causal Decoder

No SigLIP2 tower. No causal decoder. No VLM to repurpose.

Here's how it works. πŸ‘‡

(1) Raw patches, not vision-tower features Images are split into non-overlapping 32Γ—32 RGB patches and projected by a 2-layer MLP trained from scratch. Text enters through a 256-dimensional factorized embedding. Both then share every Transformer layer. β†’ 16,384-token context, enough for two 3840Γ—2160 4K UHD images

(2) Masked diffusion, not masked language modeling Pretraining is a discrete masked-diffusion text denoiser. Text-only segments draw a corruption rate from U(0,1). Multimodal segments draw from U(0.30,1), which kills the "guess it from the surrounding words" shortcut. β†’ +38.4 points masked-token accuracy from visible page patches at 90% masking (260M)

(3) Trained from scratch on a small budget About 524B packed input tokens, roughly 290B of them text-only. ModernBERT saw around 2T text tokens. They used the NorMuon optimizer to squeeze more out of the smaller budget. β†’ 16 H100s for the 260M run, 32 for the 800M

Full analysis: https://www.marktechpost.com/2026/09/06/h-company-releases-neomme-a-family-of-260m-and-800m-single-tower-multimodal-encoders-that-drop-the-vision-tower-and-causal-decoder/

Paper: https://arxiv.org/pdf/2609.01657

Technical details: https://huggingface.co/blog/Hcompany/neomme?

HF: https://huggingface.co/collections/Hcompany/neomme
  • ❀ 1
  • πŸ¦„ 1
Post #1543 809
NVIDIA Releases Personal AI Router (PAIR): An Open Inference Router That Turns The RTX, DGX Spark And Mac Boxes You Already Own Into One Local AI Cluster.

No new cluster API. No agent harness changes. No prompts leaving your network.

Here's how it works. πŸ‘‡

1. It routes, it doesn't execute
PAIR is not a new inference engine. Ollama or LM Studio still runs the model on whichever machine PAIR picks. It takes over the default port each engine uses, so the agent keeps talking to the endpoint it already knows.
β†’ Proxies Ollama-compatible, LM Studio-compatible and OpenAI-compatible endpoints
β†’ The agent decides what to request, PAIR decides where it runs

2. Discovery and trust
mDNS finds nearby machines automatically, or you add a node by IP. Trust is bootstrapped by a six-digit PIN shown on one machine and entered on the other.
β†’ All node-to-node traffic is blocked until pairing completes
β†’ Paired nodes then communicate over mTLS with generated certificates

3. The eligibility filter
A node only becomes a candidate once it can actually serve the request. The scheduler weighs five signals: is the node online and ready, is a supported engine enabled, is the exact model present, what is the current job load, what is GPU utilization.
β†’ Models don't need to be identical across nodes β€” PAIR routes by model location
β†’ Loading the same tag on more nodes just widens the eligible pool

Full analysis: https://www.marktechpost.com/2026/09/04/nvidia-releases-personal-ai-router-pair-an-open-source-virtual-inference-router-that-distributes-local-ai-requests-across-rtx-dgx-spark-and-mac-nodes/

Repo: https://github.com/NVIDIA/Personal-AI-Router

Technical details: https://www.nvidia.com/en-us/ai-on-rtx/personal-ai-router/
  • 😁 1
  • πŸ‘Œ 1
Older posts β†’

About this channel

How can I read @machinelearningresearchnews without a Telegram account?
TGViewer shows the public web preview Telegram publishes for Artificial Intelligence AI News: recent posts, photos, videos and the subscriber count, with no app, login or account.
How many subscribers does Artificial Intelligence AI News have?
Artificial Intelligence AI News (@machinelearningresearchnews) has 3.49K subscribers on Telegram, refreshed roughly every 30 minutes.
Does Artificial Intelligence AI News know I viewed it here?
No. Public channel previews carry no viewer identity, and TGViewer has no accounts or tracking of what you look up.
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook β†’Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 β†’