Trending
Trending repos, products, research papers & MCP tools
Your AI travel agent that plans, books, and adapts
a tiny buddy that lives on your Mac
The governed API layer for every app, person, and agent
Agentic content platform to create and localize content
One prompt. A whole dashboard, live in your chat.
The AI travel assistant in your pocket
The memory layer for your business
Let users control your SaaS with natural language
AI Agent that monitors your code so you're not on call
Private file tools that run on your device.
Whatever you need from government, start here.
The help center I built after Intercom's search broke
Letterboxd for music, on iPhone
The AI Agent that does 95% of your GEO work
Run coding agents on your own servers, from anywhere
Team document collab with your terminal harness (cc, codex)
Notes that appear with a mouse shake
The PHP framework your AI agent can actually inspect
a statusbar for any terminal
Start, watch and answer AI coding CLIs from one desktop app
Mute the room, keep the speaker, in real time
Say it once and your AI assistant plans the rest
An independent email API from former Postmark folks
The 2FA app you actually enjoy using
Share what your agents make, get feedback on the exact line
Accountant's Proof of the Wallet's Balance
Catch API keys the moment they appear, encrypted locally
A fast, native Jira client for Mac
Turn Any Document Into Structured Data
A smoke alarm that detects fire, not toast.
Take a car, a rocket, even your body apart in 3D
Always on agents built to handle everything
Official app for DeepSeek’s open-source agent harness
Drop the files, get a live URL in seconds
Put your MacBook notch to work: music, files, timers & more
A chatbot built on a model that can't write
Type like a human without touching the keyboard
A context aware FOSS voice assistant and dictation for macOS
Idea to physical product, engineer anything you can imagine
Open source local AI voice-to-text for Mac for private STT
an ai teleprompter that lives under your MacBook notch
One voice for all your coding agents, from your phone
Build and publish sites with your Claude or ChatGPT
Simple, private docs with smart AI templates, zero bloat.
Free, open-source Reddit and Twitter lead monitoring
Drag a box on your screen and ask AI about it
Turn notes into tasks, with AI that asks first
Share files instantly from your own cloud. Mac, Windows, iOS
Near instant multiplayer vibecoding, just point and talk
Your AI Teammate for Revenue Execution
Top 20 from alphaxiv.org · Updated Oct 1
Tokenization is presented as a core language-model design choice that shapes sequence length, computational cost, multilingual equity, evaluation, and security—not merely as preprocessing. The survey reviews established subword methods and emerging byte-level, latent, and visual alternatives, finding that BPE still dominates practice while no tokenizer consistently optimizes compression, linguistic alignment, downstream quality, and crosslingual fairness at once.
Google DeepMind announces Gemini 4 Argon, a frontier model built to sustain deep reasoning across long, multi-step workflows in software engineering, enterprise knowledge work, and cybersecurity defense. Argon is first rolling out to trusted cyber defenders through the Fairwind Program, with broader API and consumer access to follow after safeguards are strengthened. It launches at an introductory price of $2 / $10 per million input / output tokens.
Context Language Models let an agent edit its live conversation by mirroring the transcript into a file whose changes synchronize with subsequent model calls. The approach supports multiple agents and zero-shot context management. On BrowseComp-Plus, Qwen3.6-27B reached 59.4% accuracy, with 21.5% fewer prefix-reuse FLOPs than Codex-style summarization. Results vary by setting and do not isolate editing alone.
Looped-DiT improves text-to-image generation by repeating shared middle Transformer blocks within each denoising step, refining hidden representations without adding unique parameters. Deep Supervision trains intermediate predictions, and Self-Modulating Attention regulates repeated updates. Its B/16 model averages 71.5 across six benchmarks versus 69.0 for 1.7B InternVL-U, at 150 versus 732 TFLOPs per image, in 260M-parameter pixel-space models using 512 × 512 resolution.
DexAgent converts one egocentric RGB video and a task prompt into robot-training data by adapting reconstruction, motion skills, and property-specific verifiers to each task. It refines subgoals, varies simulated starting states, and renders observations for rigid, articulated, and deformable objects. Across eleven tasks, policies averaged 63.6% success versus 18.2% for SPIDER, but verification does not eliminate errors or establish causation.
T²Mem gives a vision-language-action robot policy a fixed-size episodic memory in fast weights updated during deployment, while learned policy parameters remain fixed. At each observation, a learned interface writes and reads compact vision-language associations for action selection. On RoboMME, it achieves 56.83% macro-averaged terminal success across 16 tasks, versus 17.93% for memory-free π0.5, but the result is benchmark-specific.
MatToolBench evaluates whether agents can complete materials-science workflows across specialist software, scripts, and scientific databases. It includes 204 tasks on 10 tools, with GUI, hybrid OriginPro, code, and diagnostic mixed tasks, and scores partial criteria separately from complete success. GPT-5.4 reaches a 45.0% code success rate, while Claude-sonnet-4.6 reaches 25.0% for GUI tasks; these results do not establish readiness for unsupervised use.
PRISM generates new human–object interactions from four real box-carrying videos using video generation conditioned on object categories. Contact anchors constrain reconstruction, retargeting, and policy training. The four seeds produced 256 clips, 137 feasible trajectories, and 129 successful teacher rollouts. Simulation success rose from 22.50% to 96.25% in-domain and 12.50% to 72.92% out-of-domain; box trials succeeded 15/15, but scaling remains unestablished.
PivotOPD trains multi-turn agents to avoid pivotal mistakes and recover from the states those mistakes create. A teacher supplies gold actions for prevention and recovery responses for later turns, combining these signals with group-based reinforcement learning. On ALFWorld, Qwen3-1.7B reaches 73.7% success versus 68.2% for SDAR; recovery evidence is principally oracle-labeled ALFWorld analysis.
AD-E2E-JEPA enables policy-free, goal-conditioned driving planning by compressing DINOv3 patch embeddings, predicting action-conditioned future representations, and selecting among candidate trajectories by latent distance to a future image goal. On NAVSIMv2’s 12,146-scene test set, it reaches 67.3 EPDMS at 256 candidates and 0.8 seconds per scene. More candidates improve scores but increase runtime and reduce hit rates.
LongLive-Plug distills reusable few-step sampling and guidance into adapters trained on base video models. Developers attach compatible adapters to specialized descendants without downstream retraining while retaining task-specific weights. On SCOPE, four-step Wan2.2-TI2V-5B achieves FVD 478.7, versus 805.5 for naive sampling and 502.1 for SCOPE-specific distillation on 1,378 clips. The verified scope is 54 entries, limited to compatible descendants.
Invent-a-Dataset generates training-ready text instruction and preference data from a natural-language dataset description, without requiring seed examples for the target capability. At 5,000 samples across eight datasets, it reports quality of 7.90 versus 6.76 for Claude Opus 5 and 6.73 for GLM-5.3, plus 19% higher DCScore diversity than GLM-5.3. These benchmark and Medical QA results do not establish universal gains.
ReWAM trains a robot world-action model by using an action-shaped compact visual state and a branch predicting future representations. During training, action gradients shape the state, while world-prediction gradients stop at representations; inference does not generate future states. On RoboTwin 2.0, ReWAM reaches 93.6% random-scene success versus 92.2% for AHA-WAM without generative video pre-training. The result does not establish universality.
LibraryDesignBench evaluates libraries by having one agent design a library and downstream agents use it to solve expert-validated programming problems, scoring test performance together with static-code simplicity. Across 242 problems, Opus 5.5-designed libraries scored 48.9 versus 46.6 for production libraries. This result is limited to the benchmark’s tasks, languages, consumers, and budgets.
DrivingBench tests whether an unmodified general-purpose vision-language model can steer a low-speed Toyota Corolla through a 127 m cone course using camera images and direct steering and speed commands. Four supervised model–harness systems made up to three dependent attempts; GPT-6 Astra completed it on attempt two. This preliminary result comes from one independent trial per system on a private course.
The authors study mode-hopping, abrupt changes in whether language models infer a task or follow a prompt pattern during pre-training. They use six cheap behavioral tests comparing probabilities of correct and misleading answers, plus fine-tuning tests, across OLMo3 and Apertus checkpoints. OLMo3-32B scores 81%, 0%, and 81.7% on answer+1 at 2.17T, 2.19T, and 2.21T tokens, illustrating checkpoint reversals rather than universal behavior.
SoL-Refiner refines low-resolution video drafts at target resolution in one denoising step, without changing or retraining the upstream generator. On shared-input 2K Refiner-Bench evaluation, it scores 0.81048 and 60.4150, versus 0.80351 and 56.8192 for one-step SeedVR2. These are metric averages, not human judgments; its 23-step version scores higher, and they do not establish long-range temporal consistency.
Imagine3D-LLM trains a multimodal language model to summarize views as a 3D Gaussian Splatting scene before answering questions. Summary tokens are trained to reconstruct views and predict answers, with teacher distillation only during training. On SPAR-Bench, Imagine3D-LLM-7B scores 68.5 overall versus 60.9 for a matched baseline, but the method remains limited to indoor scenes and fixed capacity.
The paper trains a proposer to revise a language-model agent’s executable harness—its workflow for model calls, tools, and information flow—using execution outcomes, while keeping the solver fixed. At test time, frozen proposer and solver weights support successive harness edits. On 21 unseen reasoning families, mean held-out score rises from 0.32 to 0.62, but five harder families are excluded.
The paper presents a meta-reasoning controller for agents reconstructing software from documentation and an execute-only reference program. The controller assesses progress, proposes and evaluates computations against the remaining model-call budget, then dispatches selected artifacts or stops. On 200-task ProgramBench, GPT-5.5 reaches 71.5% versus 63.7% for direct control at 1,200 calls, but gains do not hold at every budget.
The paper tests whether a model can generate its own training curriculum: a generator writes programs, a universal computer executes them into bytes, and a learner predicts those bytes. Adaptive self-play improves next-byte prediction on natural data without using them for gradient updates. On DCLM text, the compute exponent is 0.123. Models below 25 million parameters cannot provide contingent facts.
Alex Zhang argues that the input / output "shape" of language models has stayed fixed as the autoregressive, decoder-only Transformer since ChatGPT, so agent design has meant fitting harnesses to that shape. He proposes the reverse: designing model shapes around specific harnesses and task structures. Using Jev's [0,1]-constrained, prefill-only output and agent trajectories as examples, he argues that open models and distillation now make this direction practical for independent researchers.
Raven treats each model-plus-harness as a specialist; a host assigns work through a directed acyclic graph of dependencies. MAOB’s Exact Match Rate with Qwen3.8-27B was 0.711, versus 0.607 for Hermes Agent. Raven-Oncall achieved 82.35% success at $5.10 per case on 17 AI4S tasks; MAOB tests planning before execution, and this result does not establish causation.
The paper tests whether AI agents can alter execution traces before oversight, using local containerized harnesses, direct requests, hidden instructions, and undisclosed reward incentives. It measures attack success as trace-changing tool calls across ten trials and proposes an independent interception server with append-only logging. All ten configurations exceeded 80% on Terminal-Bench, but the experiments do not estimate real-world prevalence.
The paper tests whether language-model agents can form a covert channel while a monitor approves each update. In simulated incident response, a sender privately identifies one of four malicious connections and summarizes a public report; interaction histories and correctness feedback let a receiver learn wording without a codebook. After 60 rounds, Sol reaches 98.8% accuracy, but this is a simulated result, not evidence of deployed-system leaks.
APPL carries each skill’s structural prior into runtime selection, preserving prior-specific policies, handoff conditions, support, and verification evidence. A construction agent builds and checks policies, while a runtime agent selects one and its execution conditions. With two demonstrations per task, APPL reached 89.58% macro OOD success across six MetaWorld tasks. Results are simulated, and priors do not guarantee competence.
DCE uses an evolving gold-conditioned teacher and verified shorter rewrites to train reasoning models on their own trajectories. The latest checkpoint becomes both student and detached teacher, while accepted rewrites provide cross-entropy targets. On Qwen3-8B, DCE+SRCL reaches 65.97% Average@12 versus 30.35% for OPSD across four benchmarks. At 8B, output falls from 19,046 to 17,561 tokens relative to DCE, but rises at 1.7B.
MatToolBench evaluates whether agents can complete materials-science workflows across specialist software, scripts, and scientific databases. It includes 204 tasks on 10 tools, with GUI, hybrid OriginPro, code, and diagnostic mixed tasks, and scores partial criteria separately from complete success. GPT-5.4 reaches a 45.0% code success rate, while Claude-sonnet-4.6 reaches 25.0% for GUI tasks; these results do not establish readiness for unsupervised use.
InternW0-Δ transfers predictive video knowledge into robot control without supplying future video to the action policy at inference. Its Causal Imprint tokens summarize observed visual changes and receive future-derived training supervision, while directed attention blocks realized future frames. On RoboTwin 2.0 Clean2Random, it reaches 71.9% versus 69.4% for Qwen-RobotManip-Context; an ablation rises 70.80% to 76.45% with alignment on this benchmark.
Google DeepMind announces Gemini 4 Argon, a frontier model built to sustain deep reasoning across long, multi-step workflows in software engineering, enterprise knowledge work, and cybersecurity defense. Argon is first rolling out to trusted cyber defenders through the Fairwind Program, with broader API and consumer access to follow after safeguards are strengthened. It launches at an introductory price of $2 / $10 per million input / output tokens.
The authors study mode-hopping, abrupt changes in whether language models infer a task or follow a prompt pattern during pre-training. They use six cheap behavioral tests comparing probabilities of correct and misleading answers, plus fine-tuning tests, across OLMo3 and Apertus checkpoints. OLMo3-32B scores 81%, 0%, and 81.7% on answer+1 at 2.17T, 2.19T, and 2.21T tokens, illustrating checkpoint reversals rather than universal behavior.
The paper tests whether a Transformer can process two unrelated text streams by averaging token embeddings position by position. It compares the mixed output with the average of independent predictions and measures retained signal by rank. Across models, each stream’s independently predicted token appears in the mixed output’s top 10 in roughly 30–40% of cases, not demonstrating full distributional equality.
The authors introduce Endpoint-Constrained Optimization (ECO), which reshapes intermediate waypoints after policy inference while preserving recent executed history and the policy’s endpoint. It optimizes smoothness, turn sharpness, and fidelity without hard physical constraints. In closed-loop HUGSIM tests, VaVAM’s HD-Score rises from 18.1 to 31.0, but ECO does not guarantee collision-free driving or real-road safety.
Tokenization is presented as a core language-model design choice that shapes sequence length, computational cost, multilingual equity, evaluation, and security—not merely as preprocessing. The survey reviews established subword methods and emerging byte-level, latent, and visual alternatives, finding that BPE still dominates practice while no tokenizer consistently optimizes compression, linguistic alignment, downstream quality, and crosslingual fairness at once.
Simplex Diffusion Models (SDMs) generate sequences while retaining token uncertainty as a probability vector instead of making a categorical choice. They use closed-form transitions with churn rather than ODE integration. On TinyGSM, SDM reaches 49.0% versus 45.8% for masked diffusion at 512 evaluations; after distillation, it reaches 32.1% at 8 versus 21.4% at 128. These results come from small models and do not establish general superiority.
PISA reduces long-context sparse-attention block-selection cost by searching a hierarchy of pooled key summaries rather than scoring every fine block. It progressively retains promising regions and uses LogSumExp scores; under fixed parameters, selection complexity is O(N log N). At 256K tokens, selection took 31.440747 ms versus BSA’s 312.955505 ms, excluding attention itself; retrieval gains were evaluated only through 16K.
TrackEverything tracks dense scene points across long videos by storing surface tracks in shared 3D coordinates, merging repeated observations in spatial voxels, and decoding trajectories mainly for moving points. On the first 48 TAPVid-3D frames, it achieves 34.7 APD-P versus 10.9 for VDPM. The benchmark uses sparse queries, and memory still grows with newly encountered geometry.
The paper evaluates Jev, a model that answers multiple typed questions about one interaction state in a single call, as a detector of benchmark-labeled language-model failures. RLCDAlignBench covers 44 benchmarks, ten failure types, and 7,193 instances. A generic probability question reaches median AUROC 0.886 on 31 benchmarks, but probabilities are often miscalibrated within individual benchmarks.
Rolling-WAM carries unexecuted video-action chunks across closed-loop replanning cycles, jointly refining the future while completing only the nearest action chunk for execution. In the default 5-chunk setup, steady-state replanning takes 2 denoising steps, while each chunk still receives 10 before execution. On RoboTwin 2.0, latency was 215 ms versus 978 ms for Joint-WAM under the A100 timing setup, with competitive task success.
The authors present Rufus-Air, a reproducible eight-stage post-training recipe for GLM-4.5-Air-Base, using open data and components without new human annotation or an in-house teacher. The pipeline progresses from SFT through verifiable-reward and agent training to RLHF. It scores 76.9 on prompt-strict IFBench versus 33.6 for the same-base official release, but is not shown optimal and trails it on Creative Writing.
Top 50 from mcpmarket.com · Ranked by GitHub stars · Updated Oct 1
Empowers AI coding agents with a comprehensive and structured software development workflow, from design refinement to TDD-driven implementation.
Provides real-time global intelligence through AI-powered news aggregation, geopolitical monitoring, and infrastructure tracking in a unified situational awareness interface.
Orchestrates intelligent multi-agent swarms for Claude, coordinating autonomous workflows and building conversational AI systems.
Facilitates spec-driven development to ensure alignment between humans and AI coding assistants before any code is written.
Aggregates trending topics from over 35 platforms, offering intelligent filtering, automated multi-channel notifications, and AI-powered conversational analysis for deep news insights.
Fetches up-to-date documentation and code examples for LLMs and AI code editors directly from the source.
Empower teams with an open-source design and prototyping platform that bridges the gap between design and code through real-time collaboration.
Provides coding agents with programmatic access to Chrome DevTools for comprehensive browser control, inspection, and debugging.
Indexes codebases into a persistent knowledge graph for efficient, structural exploration and analysis by AI assistants and developers.
Create, update, and share professional resumes using a free, open-source, and privacy-focused builder.
Create secure, customizable, and privacy-focused resumes with a free and open-source builder.
Simplifies the creation, updating, and sharing of professional resumes with a focus on privacy and customizability.
Build AI applications that can learn and answer questions over large-scale federated data sources using a federated query engine.
Automates browser interactions for Large Language Models (LLMs) using Playwright.
Integrates AI capabilities with draw.io diagrams to create, modify, and enhance visualizations using natural language.
Enables advanced automation and interaction capabilities with GitHub APIs for developers and tools using the Model Context Protocol.
Automates undetected web browsing and agentic interactions with websites, bypassing captchas and blocks using an anti-detect stealth browser.
Builds and queries temporally-aware knowledge graphs tailored for AI agents operating in dynamic environments.
Provides a reliable memory layer for AI applications and AI agents, enhancing response accuracy and reducing hallucinations.
Turns an LLM into a coding agent that works directly on your codebase by providing semantic retrieval and editing capabilities.
Conducts in-depth web and local research on any topic, generating comprehensive reports with citations.
Streamline AI-driven development workflows by automating task management with Claude.
Create and manage high-performance macOS and Linux virtual machines on Apple Silicon, with built-in support for AI agents.
Connects Blender to Claude AI, enabling prompt-assisted 3D modeling, scene creation, and manipulation.
Optimizes large language model interactions by significantly reducing context window data and ensuring session continuity.
Enables open-source, local-first 3D architectural editing with a CLI, Model Context Protocol (MCP) integration, and AI agent workflows.
Autonomously builds, refines, and optimizes other AI agents using an advanced agentic coding workflow and framework knowledge base.
Connects AI assistants to n8n's vast library of workflow automation nodes, enabling AI to discover, understand, and build complex workflows efficiently.
Facilitates the creation of Model Context Protocol (MCP) servers with a Pythonic interface.
Captures screen and audio activity 24/7, transforming it into searchable memory and actionable data for AI agents.
Provides an industrial-grade, end-to-end speech recognition toolkit featuring ultra-fast transcription, multi-language support, speaker diarization, and emotion detection.
Facilitates self-hosted data scraping and no-watermark video downloads from Douyin and TikTok via an asynchronous REST API, MCP server, CLI, and web console.
Provides a lightweight memory system for coding agents through a graph-based issue tracker.
Streams free and open-source music, allowing users to search, build playlists, and listen across multiple platforms without ads or tracking.
Provides an open-source curriculum designed to teach the concepts and fundamentals of the Model Context Protocol (MCP) through practical code examples.
Automates machine learning research workflows, including idea generation, experiment execution, and iterative paper review and refinement.
Facilitates building and deploying fully-managed AI agents and long-running workflows with built-in durability and observability.
Provides a programmatic interface to automate content publishing, feed retrieval, and search functionalities on Xiaohongshu.com.
Provides AI coding agents with simplified Figma layout information via the Model Context Protocol.
Provides an Android environment within Docker containers, supporting noVNC, video recording, and a mobile cloud platform (MCP) server.
Unifies metadata for data discovery, observability, and governance through a central repository, in-depth column-level lineage, and seamless team collaboration.
Transforms any documentation website into a production-ready Claude AI skill in minutes.
Orchestrates intelligent multi-agent swarms and autonomous workflows to build advanced AI systems.
Download courses and media from over 1,800 sites, and run AI coding agents with built-in permissions and job management.
Streamlines the integration of GenAI tools with databases by handling complexities like connection pooling and authentication.
Provides a Go framework for building LLM/AI applications.
Extracts and downloads content from XiaoHongShu (RedNote), including user posts, collections, likes, albums, and search results, while removing watermarks.
Provides a reactive database backend designed for web app developers to fetch data and perform business logic with strong consistency using TypeScript.
Enhances AI coding agents by providing semantic code search and deep context from an entire codebase.
Exposes Chrome browser functionality to AI assistants for complex browser automation, content analysis, and semantic search.
Top 50 from mcpmarket.com · Ranked by GitHub stars · Updated Oct 1
Searches curated meme templates, suggests ideal joke formats, and renders custom memes in SVG, PNG, or hosted formats.
Authors, repairs, and validates custom AgentSkills and SKILL.md configurations with standardized frontmatter and bundled resources.
Configures, repairs, and verifies chat messaging channels securely using non-interactive CLI commands and secret references.
Integrates 1Password CLI to securely authenticate, manage vaults, and inject environment secrets without leaking credentials into chat or logs.
Automates inbox classification, multi-channel message routing, and asynchronous reply handling using structured task workflows.
Generates natural speech audio locally and offline using the high-performance sherpa-onnx runtime and ONNX voice models.
Delegates complex coding tasks, refactors, and issue-to-PR workflows to autonomous background CLI workers including Claude Code, Codex, and OpenCode.
Generates professional SVG, HTML, and Excalidraw diagrams for software architecture, system flows, and educational concepts.
Automates the end-to-end GitHub issue lifecycle by spawning sub-agents to implement code fixes, open pull requests, and resolve review comments.
Manages Discord operations including messaging, reactions, and channel management directly through Claude.
Enforces pre-action investigation gates before code edits, file creations, or destructive commands to prevent hallucinations and boost output quality.
Automates comprehensive verification pipelines for Quarkus applications, covering static analysis, testing, security audits, and native builds.
Analyzes raw developer prompts, identifies missing context, and crafts structured, ready-to-run prompts tailored for Claude Code and ECC workflows.
Builds automated, AI-powered data collection pipelines that scrape public web sources, enrich records with LLMs, and sync to databases.
Adapts and distributes original content across X, LinkedIn, Threads, and Bluesky while preserving authentic voice and eliminating generic clichés.
Optimizes Claude Code context windows by prompting manual compaction at logical task boundaries instead of arbitrary auto-compaction.
Integrates Jira issue tracking into coding workflows to fetch tickets, extract acceptance criteria, update statuses, and post comments.
Evaluates and visualizes whether coding agents strictly adhere to provided skills, rules, and definitions across varying prompt strictness levels.
Schedules, validates, and publishes multi-platform social media posts and media across 13 networks via SocialClaw.
Conducts multi-source web investigations using Firecrawl and Exa MCPs to synthesize comprehensive, fully cited research reports.