Historical Data

Archive

2026-09-01 以前的每日報告 169 篇,以及 98 期 AI news digest(2023 ~ 2026)

Daily Reports

169 reports before 2026-09-01
Aug 2026 31 reports
Jul 2026 31 reports
Jun 2026 30 reports
May 2026 30 reports
Apr 2026 30 reports
Mar 2026 17 reports

AI News Digests

98 issues · 原文已不保存,僅列摘要

2026

Mar 2026 1 issues
2026-03-13
not much happened today

MCP tools remain relevant for deterministic APIs despite ergonomic criticisms, with new web MCP support in Chrome v146 enabling continuous browsing agents. Persistent memory is emerging as a key differentiator for agents, with IBM improving task completion rates and multi-agent memory framed as a computer architecture challenge. Agent UX is evolving towards always-on, cross-device operation, exemplified by Perplexity Computer on iOS and Claude Code session management. Anthropic

anthropicibmperplexity-aillamaindex opus-4.6glm-5
Feb 2026 2 issues
2026-02-23
Anthropic accuses DeepSeek, Moonshot, and MiniMax of "industrial-scale distillation attacks".

Anthropic alleges industrial-scale distillation attacks on its Claude model by DeepSeek, Moonshot AI, and MiniMax, involving ~24,000 fraudulent accounts and >16M Claude exchanges to extract capabilities, raising concerns about competitive risks and safety. The community debates the difference between scraping and API-output extraction, highlighting a shift toward protecting models via API abuse resistance techniques. Meanwhile, coding agents like Codex and C

anthropicdeepseekmoonshot-aiminimax claudeclaude-3codexclaude-code
2026-02-20
not much happened today

Gemini 3.1 Pro demonstrates strong retrieval capabilities and cost efficiency compared to GPT-5.2 and Opus 4.6, though users report tooling and UI issues. The SWE-bench Verified evaluation methodology is under scrutiny for consistency, with updates bringing results closer to developer claims. Benchmarking debates arise over what frontier models truly measure, especially with ARC-AGI puzzles. Claude Opus 4.6 shows a noisy but notable 14.5-hour time horizon on software task

google-deepmindanthropiccontext-arenaartificial-analysis gemini-3.1-progpt-5.2opus-4.6sonnet-4.6

2025

Dec 2025 1 issues
2025-12-01
DeepSeek V3.2 & 3.2-Speciale: GPT5-High Open Weights, Context Management, Plans for Compute Scaling

DeepSeek launched the DeepSeek V3.2 family including Standard, Thinking, and Speciale variants with up to 131K context window and competitive benchmarks against GPT-5-High, Sonnet 4.5, and Gemini 3 Pro. The release features a novel Large Scale Agentic Task Synthesis Pipeline focusing on agentic behaviors and improvements in reinforcement learning post-training algorithms. The models are available on platforms like LM Arena with pricing around $0.28/$0.42 per

deepseek_ailm-arena deepseek-v3.2deepseek-v3.2-specialegpt-5-highsonnet-4.5
Sep 2025 1 issues
2025-09-12
not much happened today

Meta released MobileLLM-R1, a sub-1B parameter reasoning model family on Hugging Face with strong small-model math accuracy, trained on 4.2T tokens. Alibaba introduced Qwen3-Next-80B-A3B with hybrid attention, 256k context window, and improved long-horizon memory, priced competitively on Alibaba Cloud. Meta AI FAIR fixed a benchmark bug in SWE-Bench affecting agent evaluation. LiveMCP-101 benchmark shows frontier models like GPT-5 underperform on complex tasks with common

meta-ai-fairhuggingfacealibabaopenai mobilellm-r1qwen3-next-80b-a3bgpt-5
May 2025 1 issues
2025-05-30
Mary Meeker is so back: BOND Capital AI Trends report

Mary Meeker returns with a comprehensive 340-slide report on the state of AI, highlighting accelerating tech cycles, compute growth, and comparisons of ChatGPT to early Google and other iconic tech products. The report also covers enterprise traction and valuation of major AI companies. On Twitter, @tri_dao discusses an "ideal" inference architecture featuring attention variants like GTA, GLA, and DeepSeek MLA with high arithmetic intensity (~256), improving efficienc

anthropichugging-facedeepseek qwen-3-8b
Apr 2025 3 issues
2025-04-14
GPT 4.1: The New OpenAI Workhorse

OpenAI released GPT-4.1, including GPT-4.1 mini and GPT-4.1 nano, highlighting improvements in coding, instruction following, and handling long contexts up to 1 million tokens. The model achieves a 54 score on SWE-bench verified and shows a 60% improvement over GPT-4o on internal benchmarks. Pricing for GPT-4.1 nano is notably low at $0.10/1M input and $0.40/1M output. GPT-4.5 Preview is being deprecated in favor of GPT-4.1. Integration

openaillama-indexperplexity-aigoogle-deepmind gpt-4.1gpt-4.1-minigpt-4.1-nanogpt-4o
2025-04-07
Llama 4's Controversial Weekend Release

Meta released Llama 4, featuring two new medium-size MoE open models and a promised 2 Trillion parameter "behemoth" model, aiming to be the largest open model ever. The release included advanced training techniques like Chameleon-like early fusion with MetaCLIP, interleaved chunked attention without RoPE, native FP8 training, and training on up to 40 trillion tokens. Despite the hype, the release faced criticism for lack of transparency compared to Llama 3, implementation issues, and poo

meta llama-4llama-3llama-3-2
2025-04-03
not much happened today

Gemini 2.5 Pro shows strengths and weaknesses, notably lacking LaTex math rendering unlike ChatGPT, and scored 24.4% on the 2025 US AMO. DeepSeek V3 ranks 8th and 12th on recent leaderboards. Qwen 2.5 models have been integrated into the PocketPal app. Research from Anthropic reveals that Chains-of-Thought (CoT) reasoning is often unfaithful, especially on harder tasks, raising safety concerns. OpenAI's PaperBench benchmark shows AI agents struggle wit

googleanthropicopenaillama_index gemini-2.5-prochatgptdeepseek-v3qwen-2.5
Mar 2025 8 issues
2025-03-31
>$41B raised today (OpenAI @ 300b, Cursor @ 9.5b, Etched @ 1.5b)

OpenAI is preparing to release a highly capable open language model, their first since GPT-2, with a focus on reasoning and community feedback, as shared by @kevinweil and @sama. DeepSeek V3 0324 has achieved the #5 spot on the Arena leaderboard, becoming the top open model with an MIT license and cost advantages. Gemini 2.5 Pro is noted for outperforming models like Claude 3.7 Sonnet in coding tasks, with upcoming pricing and improvements expected soon. New startups like

openaideepseekgeminicursor deepseek-v3-0324gemini-2.5-proclaude-3.7-sonnet
2025-03-24
Halfmoon is Reve Image: a new SOTA Image Model from ex-Adobe/Stability trio

Reve, a new composite AI model from former Adobe and Stability alums Christian Cantrell, Taesung Park, and Michaël Gharbi, has emerged as the top-rated image generation model, surpassing previous state-of-the-art models like Recraft and Ideogram in text rendering and typography. The team emphasizes "enhancing visual generative models with logic" and "understanding user intent with advanced language capabilities" to iteratively amend visuals based on natural language input. Ad

artificial-analysisstability-aiadobedeepseek deepseek-v3-0324qwen-2.5-vl-32b-instructrecraft
2025-03-21
lots of little things happened this week

Anthropic introduced a novel 'think' tool enhancing instruction adherence and multi-step problem solving in agents, with combined reasoning and tool use demonstrated by Claude. NVIDIA's Llama-3.3-Nemotron-Super-49B-v1 ranked #14 on LMArena, noted for strong math reasoning and a 15M post-training dataset. Sakana AI launched a Sudoku-based reasoning benchmark to advance AI problem-solving capabilities. Meta AI released SWEET-RL, a reinforcement learning algorithm improv

anthropicnvidiasakana-aimeta-ai-fair llama-3-3-nemotron-super-49b-v1claude
2025-03-19
Every 7 Months: The Moore's Law for Agent Autonomy

METR published a paper measuring AI agent autonomy progress, showing it has doubled every 7 months since 2019 (GPT-2). They introduced a new metric, the 50%-task-completion time horizon, where models like Claude 3.7 Sonnet achieve 50% success in about 50 minutes. Projections estimate 1 day autonomy by 2028 and 1 month autonomy by late 2029. Meanwhile, Nvidia released Cosmos-Transfer1 for conditional world generation and GR00T-N1-2B, an open foundation model fo

metrnvidiahugging-facecanopy-labs claude-3-7-sonnetllama-4phi-4-multimodalgpt-2
2025-03-17
Cohere's Command A claims #3 open model spot (after DeepSeek and Gemma)

Cohere's Command A model has solidified its position on the LMArena leaderboard, featuring an open-weight 111B parameter model with an unusually long 256K context window and competitive pricing. Mistral AI released the lightweight, multilingual, and multimodal Mistral AI Small 3.1 model, optimized for single RTX 4090 or Mac 32GB RAM setups, with strong performance on instruct and multimodal benchmarks. The new OCR model SmolDocling offers fast document reading with low VR

coheremistral-aihugging-face command-amistral-ai-small-3.1smoldoclingqwen-2.5-vl
2025-03-11
The new OpenAI Agents Platform

OpenAI introduced a comprehensive suite of new tools for AI agents, including the Responses API, Web Search Tool, Computer Use Tool, File Search Tool, and an open-source Agents SDK with integrated observability tools, marking a significant step towards the "Year of Agents." Meanwhile, Reka AI open-sourced Reka Flash 3, a 21B parameter reasoning model that outperforms o1-mini and powers their Nexus platform, with weights available on Hugging Face. The *

openaireka-aihugging-facedeepseek reka-flash-3o1-miniclaude-3-7-sonnetllama-3-3-70b
2025-03-06
not much happened today

AI21 Labs launched Jamba 1.6, touted as the best open model for private enterprise deployment, outperforming Cohere, Mistral, and Llama on benchmarks like Arena Hard. Mistral AI released a state-of-the-art multimodal OCR model with multilingual and structured output capabilities, available for on-prem deployment. Alibaba Qwen introduced QwQ-32B, an open-weight reasoning model with 32B parameters and cost-effective usage, showing competitive benchmark scores. *

ai21-labsmistral-aialibabaopenai jamba-1.6mistral-ocrqwq-32bo1
2025-03-03
Anthropic's $61.5B Series E

Anthropic raised a $3.5 billion Series E funding round at a $61.5 billion valuation, signaling strong financial backing for the Claude AI model. GPT-4.5 achieved #1 rank across all categories on the LMArena leaderboard, excelling in multi-turn conversations, coding, math, creative writing, and style control. DeepSeek R1 tied with GPT-4.5 for top performance on hard prompts with style control. Discussions highlighted comparisons between GPT-4.5 and Claude 3.7 Son

anthropicopenaideepseeklmsys gpt-4.5claude-3.7-sonnetdeepseek-r1
Feb 2025 5 issues
2025-02-24
Claude 3.7 Sonnet

Anthropic launched Claude 3.7 Sonnet, their most intelligent model to date featuring hybrid reasoning with two thinking modes: near-instant and extended step-by-step thinking. The release includes Claude Code, an agentic coding tool in limited preview, and supports a 128k output token capability in beta. Claude 3.7 Sonnet performs well on coding benchmarks like SWE-Bench Verified and Cognition's junior-dev eval, and introduces advanced features such as streaming thinking,

anthropic claude-3-7-sonnetclaude-3claude-code
2025-02-19
The Ultra-Scale Playbook: Training LLMs on GPU Clusters

Huggingface released "The Ultra-Scale Playbook: Training LLMs on GPU Clusters," an interactive blogpost based on 4000 scaling experiments on up to 512 GPUs, providing detailed insights into modern GPU training strategies. DeepSeek introduced the Native Sparse Attention (NSA) model, gaining significant community attention, while Perplexity AI launched R1-1776, an uncensored and unbiased version of DeepSeek's R1 model. Google DeepMind unveiled PaliGemma 2 Mix, a multi-task visi

huggingfacedeepseekperplexity-aigoogle-deepmind deepseek-native-sparse-attentionr1-1776paligemma-2-mixmuse
2025-02-17
LLaDA: Large Language Diffusion Models

LLaDA (Large Language Diffusion Model) 8B is a breakthrough diffusion-based language model that rivals LLaMA 3 8B while training on 7x fewer tokens (2 trillion tokens) and using 0.13 million H800 GPU hours. It introduces a novel text generation approach by predicting uniformly masked tokens in a diffusion process, enabling multi-turn dialogue and instruction-following. Alongside, StepFun AI released two major models: Step-Video-T2V 30B, a text-to-video model generating up

stepfun-aiscale-aicambridgellamaindex llada-8bllama-3-8bstep-video-t2v-30bstep-audio-chat-132b
2025-02-10
not much happened today

Google released Gemini 2.0 Flash Thinking Experimental 1-21, a vision-language reasoning model with a 1 million-token context window and improved accuracy on science, math, and multimedia benchmarks, surpassing DeepSeek-R1 but trailing OpenAI's o1. ZyphraAI launched Zonos, a multilingual Text-to-Speech model with instant voice cloning and controls for speaking rate, pitch, and emotions, running at ~2x real-time speed on RTX 4090. Hugging Face released

googlezyphraaihugging-faceanthropic gemini-2.0-flash-thinking-experimental-1-21zonosopenr1-math-220khuginn-3.5b
2025-02-03
OpenAI takes on Gemini's Deep Research

OpenAI released the full version of the o3 agent, with a new Deep Research variant showing significant improvements on the HLE benchmark and achieving SOTA results on GAIA. The release includes an "inference time scaling" chart demonstrating rigorous research, though some criticism arose over public test set results. The agent is noted as "extremely simple" and currently limited to 100 queries/month, with plans for a higher-rate version. Reception has been mostly positive, wi

openaigoogle-deepmindnyuuc-berkeley o3o3-mini-higho3-deep-research-mini
Jan 2025 5 issues
2025-01-27
DeepSeek #1 on US App Store, Nvidia stock tanks -17%

DeepSeek has made a significant cultural impact by hitting mainstream news unexpectedly in 2025. The DeepSeek-R1 model features a massive 671B parameter MoE architecture and demonstrates chain-of-thought (CoT) capabilities comparable to OpenAI's o1 at a lower cost. The DeepSeek V3 model trains a 236B parameter model 42% faster than its predecessor using fp8 precision. The Qwen2.5 multimodal models support images and videos with sizes ranging from 3B to 72B p

deepseekopenainvidialangchain deepseek-r1deepseek-v3qwen2.5-vlo1
2025-01-20
DeepSeek R1: o1-level open weights model and a simple recipe for upgrading 1.5B models to Sonnet/4o level

DeepSeek released DeepSeek R1, a significant upgrade over DeepSeek V3 from just three weeks prior, featuring 8 models including full-size 671B MoE models and multiple distillations from Qwen 2.5 and Llama 3.1/3.3. The models are MIT licensed, allowing finetuning and distillation. Pricing is notably cheaper than o1 by 27x-50x. The training process used GRPO (reward for correctness and style outcomes) without relying on PRM, MCTS, or reward models, focusing on reasoning

deepseekollamaqwenllama deepseek-r1deepseek-v3qwen-2.5llama-3.1
2025-01-13
not much happened today

Helium-1 Preview by kyutai_labs is a 2B-parameter multilingual base LLM outperforming Qwen 2.5, trained on 2.5T tokens with a 4096 context size using token-level distillation from a 7B model. Phi-4 (4-bit) was released in lmstudio on an M4 max, noted for speed and performance. Sky-T1-32B-Preview is a $450 open-source reasoning model matching o1's performance with strong benchmark scores. Codestral 25.01 by mistralai is a new SOTA coding

kyutai-labslmstudiomistralaillamaindex helium-1qwen-2.5phi-4sky-t1-32b-preview
2025-01-06
PRIME: Process Reinforcement through Implicit Rewards

Implicit Process Reward Models (PRIME) have been highlighted as a significant advancement in online reinforcement learning, trained on a 7B model with impressive results compared to gpt-4o. The approach builds on the importance of process reward models established by "Let's Verify Step By Step." Additionally, AI Twitter discussions cover topics such as proto-AGI capabilities with claude-3.5-sonnet, the role of compute scaling for Artificial Superintelligence (ASI), an

openaitogether-aideepseeklangchain claude-3.5-sonnetgpt-4odeepseek-v3gemini-2.0
2025-01-03
not much happened today

Olmo 2 released a detailed tech report showcasing full pre, mid, and post-training details for a frontier fully open model. PRIME, an open-source reasoning solution, achieved 26.7% pass@1, surpassing GPT-4o in benchmarks. Performance improvements include Qwen 32B (4-bit) generating at >40 tokens/sec on an M4 Max and libvips being 25x faster than Pillow for image resizing. New tools like Swaggo/swag for Swagger 2.0 documentation, Jujutsu (jj) Git-co

olmoopenaiqwencerebras-systems primegpt-4oqwen-32b

2024

Dec 2024 7 issues
2024-12-30
not much happened today

Sam Altman publicly criticizes DeepSeek and Qwen models, sparking debate about OpenAI's innovation claims and reliance on foundational research like the Transformer architecture. Deepseek V3 shows significant overfitting issues in the Misguided Attention evaluation, solving only 22% of test prompts, raising concerns about its reasoning and finetuning. Despite skepticism about its open-source status, Deepseek V3 is claimed to surpass ChatGPT4 as an open-sou

openaideepseekgoogleqwen deepseek-v3chatgpt-4
2024-12-26
DeepSeek v3: 671B finegrained MoE trained for $5.5m USD of compute on 15T tokens

DeepSeek-V3 has launched with 671B MoE parameters and trained on 14.8T tokens, outperforming GPT-4o and Claude-3.5-sonnet in benchmarks. It was trained with only 2.788M H800 GPU hours, significantly less than Llama-3's 30.8M GPU-hours, showcasing major compute efficiency and cost reduction. The model is open-source and deployed via Hugging Face with API support. Innovations include native FP8 mixed precision training, Multi-Head Latent Attention scaling, disti

deepseek-aihugging-faceopenaianthropic deepseek-v3gpt-4oclaude-3.5-sonnetllama-3
2024-12-23
not much happened this weekend

o3 model gains significant attention with discussions around its capabilities and implications, including an OpenAI board member referencing "AGI." LangChain released their State of AI 2024 survey. Hume announced OCTAVE, a 3B parameter API-only speech-language model with voice cloning. x.ai secured a $6B Series C funding round. Discussions highlight inference-time scaling, model ensembles, and the surprising generalization ability of small models. New

openailangchainhumex-ai o3o1opussonnet
2024-12-16
Meta Apollo - Video Understanding up to 1 hour, SOTA Open Weights

Meta released Apollo, a new family of state-of-the-art video-language models available in 1B, 3B, and 7B sizes, featuring "Scaling Consistency" for efficient scaling and introducing ApolloBench, which speeds up video understanding evaluation by 41× across five temporal perception categories. Google Deepmind launched Veo 2, a 4K video generation model with improved physics and camera control, alongside an enhanced Imagen 3 image model. OpenAI globally rolled ou

meta-ai-fairhugging-facegoogle-deepmindopenai apollo-1bapollo-3bapollo-7bveo-2
2024-12-13
Meta BLT: Tokenizer-free, Byte-level LLM

Meta AI introduces the Byte Latent Transformer (BLT), a tokenizer-free architecture that dynamically forms byte patches for efficient compute allocation, outperforming Llama 3 on benchmarks including the CUTE benchmark. The model was trained on approximately 1 trillion tokens and features a three-block transformer design with local and global components. This approach challenges traditional tokenization and may enable new multimodal capabilities such as direct file interaction wi

meta-ai-fairllamaindexmicrosoftdeepseek-ai byte-latent-transformerllama-3phi-4gpt-4o
2024-12-09
OpenAI Sora Turbo and Sora.com

OpenAI launched Sora Turbo, enabling text-to-video generation for ChatGPT Plus and Pro users with monthly generation limits and regional restrictions in Europe and the UK. Google announced a quantum computing breakthrough with the development of the Willow chip, potentially enabling commercial quantum applications. Discussions on O1 model performance highlighted its lag behind Claude 3.5 Sonnet and Gemini in coding tasks, with calls for algorithmic innovation beyond t

openaigooglenvidiahugging-face sora-turboo1claude-3.5-sonnetclaude-3.5
2024-12-03
Olympus has dropped (aka, Amazon Nova Micro|Lite|Pro|Premier|Canvas|Reel)

Amazon announced the Amazon Nova family of multimodal foundation models at AWS Re:Invent, available immediately with no waitlist in configurations like Micro, Lite, Pro, Canvas, and Reel, with Premier and speech-to-speech coming next year. These models offer 2-4x faster token speeds and are 25%-400% cheaper than competitors like Anthropic Claude models, positioning Nova as a serious contender in AI engineering. Pricing undercuts models such as Google DeepMind Gemini Flash 8

amazonanthropicgoogle-deepmindsakana-ai-labs amazon-novaclaude-3llama-3-70bgemini-1.5-flash
Nov 2024 4 issues
2024-11-25
Anthropic launches the Model Context Protocol

Anthropic has launched the Model Context Protocol (MCP), an open protocol designed to enable seamless integration between large language model applications and external data sources and tools. MCP supports diverse resources such as file contents, database records, API responses, live system data, screenshots, and logs, identified by unique URIs. It also includes reusable prompt templates, system and API tools, and JSON-RPC 2.0 transports with streaming support. MCP allows servers to requ

anthropicamazonzedsourcegraph claude-3.5-sonnetclaude-desktop
2024-11-18
Pixtral Large (124B) beats Llama 3.2 90B with updated Mistral Large 24.11

Mistral has updated its Pixtral Large vision encoder to 1B parameters and released an update to the 123B parameter Mistral Large 24.11 model, though the update lacks major new features. Pixtral Large outperforms Llama 3.2 90B on multimodal benchmarks despite having a smaller vision adapter. Mistral's Le Chat chatbot received comprehensive feature updates, reflecting a company focus on product and research balance as noted by Arthur Mensch. SambaNova sponsors infer

mistral-aisambanovanvidia pixtral-largemistral-large-24.11llama-3-2qwen2.5-7b-instruct-abliterated-v2-gguf
2024-11-11
FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Epoch AI collaborated with over 60 leading mathematicians to create the FrontierMath benchmark, a fresh set of hundreds of original math problems with easy-to-verify answers, aiming to challenge current AI models. The benchmark reveals that all tested models, including o1, perform poorly, highlighting the difficulty of complex problem-solving and Moravec's paradox in AI. Key AI developments include the introduction of Mixture-of-Transformers (MoT), a sparse multi-modal tr

epoch-aiopenaimicrosoftanthropic o1claude-3.5-haikugpt-4o
2024-11-04
OpenAI beats Anthropic to releasing Speculative Decoding

Prompt lookup and Speculative Decoding techniques are gaining traction with implementations from Cursor, Fireworks, and teased features from Anthropic. OpenAI has introduced faster response times and file edits with these methods, offering about 50% efficiency improvements. The community is actively exploring AI engineering use cases with these advancements. Recent updates highlight progress from companies like NVIDIA, OpenAI, Anthropic, Microsoft, B

openaianthropicnvidiamicrosoft claude-3-sonnetmrt5
Oct 2024 5 issues
2024-10-29
GitHub Copilot Strikes Back

GitHub's tenth annual Universe conference introduced the Multi-model Copilot featuring Anthropic's Claude 3.5 Sonnet, Google's Gemini 1.5 Pro, and OpenAI's o1-preview models in a new picker UI, allowing developers to choose from multiple companies' models. The event also showcased GitHub Spark, an AI-native tool for building natural language applications with deployment-free hosting and integrated model prompting. Additionally, GitHub updated its Copilot Workspace with ne

githubanthropicgoogle-deepmindopenai claude-3-5-sonnetgemini-1.5-proo1-previewgemini-flash-8b
2024-10-21
DocETL: Agentic Query Rewriting and Evaluation for Complex Document Processing

UC Berkeley's EPIC lab introduces innovative LLM data operators with projects like LOTUS and DocETL, focusing on effective programming and computation over large data corpora. This approach contrasts GPU-rich big labs like Deepmind and OpenAI with GPU-poor compound AI systems. Microsoft open-sourced BitNet b1.58, a 1-bit ternary parameter LLM enabling 4-20x faster training and on-device inference at human reading speeds. Nvidia released Llama-3.1-Nemotron-70B-In

uc-berkeleydeepmindopenaimicrosoft bitnet-b1.58llama-3.1-nemotron-70b-instructgpt-4oclaude-3.5-sonnet
2024-10-16
Did Nvidia's Nemotron 70B train on test?

NVIDIA's Nemotron-70B model has drawn scrutiny despite strong benchmark performances on Arena Hard, AlpacaEval, and MT-Bench, with some standard benchmarks like GPQA and MMLU Pro showing no improvement over the base Llama-3.1-70B. The new HelpSteer2-Preference dataset improves some benchmarks with minimal losses elsewhere. Meanwhile, Mistral released Ministral 3B and 8B models featuring 128k context length and outperforming Llama-3.1 and GPT-4o

nvidiamistral-aihugging-facezep nemotron-70bllama-3.1-70bllama-3.1ministral-3b
2024-10-07
not much happened this weekend

AI news from 10/4/2024 to 10/7/2024 highlights several developments: OpenAI's o1-preview shows strong performance on complex tasks but struggles with simpler ones, while Claude 3.5 Sonnet can match its reasoning through advanced prompting techniques. Meta introduced Movie Gen, a cutting-edge media foundation model for text-to-video generation and editing. Reka updated their 21B Flash Model with temporal video understanding, native audio, and tool use capabilities. Interes

openaimeta-ai-fairrekalangchainai o1-previewclaude-3.5-sonnet21b-flash-model
2024-10-04
Contextual Document Embeddings: `cde-small-v1`

Meta announced a new text-to-video model, Movie Gen, claiming superior adaptation of Llama 3 to video generation compared to OpenAI's Sora Diffusion Transformers, though no release is available yet. Researchers Jack Morris and Sasha Rush introduced the cde-small-v1 model with a novel contextual batching training technique and contextual embeddings, achieving strong performance with only 143M parameters. OpenAI launched Canvas, a collaborative interface for ChatGPT

meta-ai-fairopenaigoogle-deepmindweights-biases llama-3cde-small-v1gemini-1.5-flash-8bchatgpt
Sep 2024 6 issues
2024-09-30
Liquid Foundation Models: A New Transformers alternative + AINews Pod 2

Liquid.ai emerged from stealth with three subquadratic foundation models demonstrating superior efficiency compared to state space models and Apple’s on-device and server models, backed by a $37M seed round. Meta AI announced Llama 3.2 with multimodal vision-enabled models and lightweight text-only variants for mobile. Google DeepMind introduced production-ready Gemini-1.5-Pro-002 and Gemini-1.5-Flash-002 models with improved pricing and rate limits, alongside AlphaChip

liquid-aimeta-ai-fairgoogle-deepmindopenai llama-3-2gemini-1.5-pro-002gemini-1.5-flash-002
2024-09-24
ChatGPT Advanced Voice Mode

OpenAI rolled out ChatGPT Advanced Voice Mode with 5 new voices and improved accent and language support, available widely in the US. Ahead of rumored updates for Llama 3 and Claude 3.5, Gemini Pro saw a significant price cut aligning with the new intelligence frontier pricing. OpenAI's o1-preview model showed promising planning task performance with 52.8% accuracy on Randomized Mystery Blocksworld. Anthropic is rumored to release a new model, generating community exc

openaianthropicscale-aitogethercompute o1-previewqwen-2.5llama-3claude-3.5
2024-09-19
not much happened today

OpenAI's o1-preview and o1-mini models lead benchmarks in Math, Hard Prompts, and Coding. Qwen 2.5 72B model shows strong performance close to GPT-4o. DeepSeek-V2.5 tops Chinese LLMs, rivaling GPT-4-Turbo-2024-04-09. Microsoft's GRIN MoE achieves good results with 6.6B active parameters. Moshi voice model from Kyutai Labs runs locally on Apple Silicon Macs. Perplexity app introduces voice mode with push-to-talk. LlamaCoder by Together.ai uses Llama 3.1 405B*

openaiqwendeepseek-aimicrosoft o1-previewo1-miniqwen-2.5gpt-4o
2024-09-16
a quiet weekend

OpenAI released the new o1 model, leveraging reinforcement learning and chain-of-thought prompting to excel in reasoning benchmarks, achieving an IQ-like score of 120. Google DeepMind introduced DataGemma to reduce hallucinations by connecting LLMs with real-world data, and unveiled ALOHA and DemoStart for robot dexterity using diffusion methods. Adobe previewed its Firefly AI Video Model with text-to-video and generative extend features. Mistral launched

openaigoogle-deepmindadobemistral-ai o1datagemmaalohademostart
2024-09-10
not much happened today + AINews Podcast?

Glean doubled its valuation again. Dan Hendrycks' Superforecaster AI generates plausible election forecasts with interesting prompt engineering. A Stanford study found that LLM-generated research ideas are statistically more novel than those by expert humans. SambaNova announced faster inference for llama-3 models, surpassing Cerebras. Benjamin Clavie gave a notable talk on retrieval-augmented generation techniques. Strawberry is reported to launch in two week

gleansambanovacerebrasstanford superforecaster-aillama-3reflection-70b
2024-09-03
Everybody shipped small things this holiday weekend

xAI announced the Colossus 100k H100 cluster capable of training an FP8 GPT-4 class model in 4 days. Google introduced Structured Output for Gemini. Anthropic discussed Claude's performance issues possibly due to API prompt modifications. OpenAI enhanced controls for File Search in their Assistants API. Cognition and Anthropic leaders appeared on podcasts. The viral Kwai-Kolors virtual try-on model and the open-source real-time audio conversational mod

xaigoogleanthropicopenai gpt-4o-voicegeminiclaudejamba-1.5
Aug 2024 4 issues
2024-08-26
not much happened this weekend

Nous Research announced DisTrO, a new optimizer that drastically reduces inter-GPU communication by 1000x to 10,000x enabling efficient training on slow networks, offering an alternative to GDM's DiLoCo. Cursor AI gained viral attention from an 8-year-old user and announced a new fundraise, with co-host Aman returning to their podcast. George Hotz launched tinybox for sale. In robotics, AGIBOT revealed 5 new humanoid robots with open-source plans, and Unitree show

nous-researchcursor-aigdmgeorge-hotz jamba-1.5dream-machine-1.5ideogram-v2mistral-nemo-minitron-8b
2024-08-19
The DSPy Roadmap

Omar Khattab announced joining Databricks before his MIT professorship and outlined the roadmap for DSPy 2.5 and 3.0+, focusing on improving core components like LMs, signatures, optimizers, and assertions with features such as adopting LiteLLM to reduce code and enhance caching and streaming. The roadmap also includes developing more accurate, cost-effective optimizers, building tutorials, and enabling interactive optimization tracking. On AI Twitter, Google launched Gemin

databricksmitgoogleopenai dspylitel-lmgeminichatgpt-4o
2024-08-13
Gemini Live

Google launched Gemini Live on Android for Gemini Advanced subscribers during the Pixel 9 event, featuring integrations with Google Workspace apps and other Google services. The rollout began on 8/12/2024, with iOS support planned. Anthropic released Genie, an AI software engineering system achieving a 57% improvement on SWE-Bench. TII introduced Falcon Mamba, a 7B attention-free open-access model scalable to long sequences. Benchmarking showed that longer context

googleanthropictiisupabase gemini-1.5-progeniefalcon-mambagemini-1.5
2024-08-06
GPT4o August + 100% Structured Outputs for All (GPT4o mini edition)

Stability.ai users are leveraging LoRA and ControlNet for enhanced line art and artistic style transformations, while facing challenges with AMD GPUs due to the discontinuation of ZLUDA. Community tensions persist around the r/stablediffusion subreddit moderation. Unsloth AI users report fine-tuning difficulties with LLaMA3 models, especially with PPO trainer integration and prompt formatting, alongside anticipation for multi-GPU support and cost-effective clo

stability-aiunsloth-aigooglehugging-face gpt-4o-minigpt-4o-2024-08-06llama-3bigllama-3.1-1t-instruct
Jul 2024 8 issues
2024-07-29
Apple Intelligence Beta + Segment Anything Model 2

Meta advanced its open source AI with a sequel to the Segment Anything Model, enhancing image segmentation with memory attention for video applications using minimal data and compute. Apple Intelligence delayed its official release to iOS 18.1 in October but launched developer previews on MacOS Sequoia, iOS 18, and iPadOS 18, accompanied by a detailed 47-page paper revealing extensive pretraining on 6.3T tokens and use of Cloud TPUs rather than Apple Silicon. The

meta-ai-fairapple llama-3-405bllama-3segment-anything-model
2024-07-25
AlphaProof + AlphaGeometry2 reach 1 point short of IMO Gold

Search+Verifier highlights advances in neurosymbolic AI during the 2024 Math Olympics. Google DeepMind's combination of AlphaProof and AlphaGeometry 2 solved four out of six IMO problems, with AlphaProof being a finetuned Gemini model using an AlphaZero approach, and AlphaGeometry 2 trained on significantly more synthetic data with a novel knowledge-sharing mechanism. Despite impressive results, human judges noted the AI required much longer time than human competitors. Meanw

google-deepmindmeta-ai-fairmistral-ai geminialphageometry-2alphaproofllama-3-1-405b
2024-07-22
Llama 3.1 Leaks: big bumps to 8B, minor bumps to 70b, and SOTA OSS 405b model

Llama 3.1 leaks reveal a 405B dense model with 128k context length, trained on 39.3M GPU hours using H100-80GB GPUs, and fine-tuned with over 25M synthetic examples. The model shows significant benchmark improvements, especially for the 8B and 70B variants, with some evals suggesting the 70B outperforms GPT-4o. GPT-4o Mini launched as a cost-efficient variant with strong performance but some reasoning weaknesses. Synthetic datasets like NuminaMath enable models su

meta-ai-fairopenaialibaba llama-3-1-405bllama-3-8bllama-3-70bllama-3-1-8b
2024-07-18
Mini, Nemo, Turbo, Lite - Smol models go brrr (GPT4o version)

GPT-4o-mini launches with a 99% price reduction compared to text-davinci-003, offering 3.5% the price of GPT-4o and matching Opus-level benchmarks. It supports 16k output tokens, is faster than previous models, and will soon support text, image, video, and audio inputs and outputs. Mistral Nemo, a 12B parameter model developed with Nvidia, features a 128k token context window, FP8 checkpoint, and strong benchmark performance. Together Lite and Turbo offer

openainvidiamistral-aitogethercompute gpt-4o-minimistral-nemollama-3llama-3-400b
2024-07-15
Microsoft AgentInstruct + Orca 3

Microsoft Research released AgentInstruct, the third paper in its Orca series, introducing a generative teaching pipeline that produces 25.8 million synthetic instructions to fine-tune mistral-7b, achieving significant performance gains: +40% AGIEval, +19% MMLU, +54% GSM8K, +38% BBH, +45% AlpacaEval, and a 31.34% reduction in hallucinations. This synthetic data approach follows the success of FineWeb and Apple's Rephrasing research in improving dataset quality. Additi

microsoft-researchappletencenthugging-face mistral-7borca-2.5
2024-07-08
Problems with MMLU-Pro

MMLU-Pro is gaining attention as the successor to MMLU on the Open LLM Leaderboard V2 by HuggingFace, despite community concerns about evaluation discrepancies and prompt sensitivity affecting model performance, notably a 10-point improvement in Llama-3-8b-q8 with simple prompt tweaks. Meta's MobileLLM research explores running sub-billion parameter LLMs on smartphones using shared weights and deeper architectures. Salesforce's APIGen introduces an automated dataset g

huggingfacemeta-ai-fairsalesforcerunway mmlu-prollama-3-8b-q8gpt4all-3.0chatgpt
2024-07-05
Qdrant's BM42: "Please don't trust us"

Qdrant attempted to replace BM25 and SPLADE with a new method called "BM42" combining transformer attention and collection-wide statistics for semantic and keyword search, but their evaluation using the Quora dataset was flawed. Nils Reimers from Cohere reran BM42 on better datasets and found it underperformed. Qdrant acknowledged the errors but still ran a suboptimal BM25 implementation. This highlights the importance of dataset choice and evaluation sanity checks in search model cl

qdrantcoherestripeanthropic claude-3.5-sonnetgemma-2nano-llava-1.5
2024-07-01
RouteLLM: RIP Martian? (Plus: AINews Structured Summaries update)

LMSys introduces RouteLLM, an open-source router framework trained on preference data from Chatbot Arena, achieving cost reductions over 85% on MT Bench, 45% on MMLU, and 35% on GSM8K while maintaining 95% of GPT-4's performance. This approach surpasses previous task-specific routing by using syntax-based Mixture of Experts (MoE) routing and data augmentation, beating commercial solutions by 40%. The update highlights advances in LLM routing, cost-efficiency, and model

lmsysopenai gpt-4gemma-2-27bgemma-2-9b
Jun 2024 4 issues
2024-06-19
There's Ilya!

Ilya Sutskever has co-founded Safe Superintelligence Inc shortly after leaving OpenAI, while Jan Leike moved to Anthropic. Meta released new models including Chameleon 7B and 34B with mixed-modal input and unified token space quantization. DeepSeek-Coder-V2 shows code capabilities comparable to GPT-4 Turbo, supporting 338 programming languages and 128K context length. Consistency Large Language Models (CLLMs) enable parallel decoding generating

safe-superintelligence-incopenaianthropicmeta chameleon-7bchameleon-34bdeepseek-coder-v2gpt-4-turbo
2024-06-17
Is this... OpenQ*?

DeepSeekCoder V2 promises GPT4T-beating performance at a fraction of the cost. Anthropic released new research on reward tampering. Runway launched their Sora response and Gen-3 Alpha video generation model. A series of papers explore "test-time" search techniques improving mathematical reasoning with models like LLaMa-3 8B. Apple announced Apple Intelligence with smarter Siri and image/document understanding, partnered with OpenAI to integrate ChatGPT into iOS 18, and re

deepseek_aianthropicrunwaymlopenai deepseek-coder-v2llama-3-8bnemotron-4-340bstable-diffusion-3-medium
2024-06-10
Talaria: Apple's new MLOps Superweapon

Apple Intelligence introduces a small (~3B parameters) on-device model and a larger server model running on Apple Silicon with Private Cloud Compute, aiming to surpass Google Gemma, Mistral Mixtral, Microsoft Phi, and Mosaic DBRX. The on-device model features a novel lossless quantization strategy using mixed 2-bit and 4-bit LoRA adapters averaging 3.5 bits-per-weight, enabling dynamic adapter hot-swapping and efficient memory management. Apple credits the Talaria tool fo

applegooglemistral-aimicrosoft gemmamixtralphidbrx
2024-06-05
5 small news items

OpenAI announces that ChatGPT's voice mode is "coming soon." Leopold Aschenbrenner launched a 5-part AGI timelines series predicting a trillion dollar cluster from current AI progress. Will Brown released a comprehensive GenAI Handbook. Cohere completed a $450 million funding round at a $5 billion valuation. DeepMind research on uncertainty quantification in LLMs and an xLSTM model outperforming transformers were highlighted. Studies on the geometry of conce

openaicoheredeepmindhugging-face llama-3xLSTM
May 2024 5 issues
2024-05-30
Contextual Position Encoding (CoPE)

Meta AI researcher Jason Weston introduced CoPE, a novel positional encoding method for transformers that incorporates context to create learnable gates, enabling improved handling of counting and copying tasks and better performance on language modeling and coding. The approach can potentially be extended with external memory for gate calculation. Google DeepMind released Gemini 1.5 Flash and Pro models optimized for fast inference. Anthropic announced general avai

meta-ai-fairgoogle-deepmindanthropicperplexity-ai copegemini-1.5-flashgemini-1.5-proclaude
2024-05-27
Life after DPO (RewardBench)

xAI raised $6 billion at a $24 billion valuation, positioning it among the most highly valued AI startups, with expectations to fund GPT-5 and GPT-6 class models. The RewardBench tool, developed by Nathan Lambert, evaluates reward models (RMs) for language models, showing Cohere's RMs outperforming open-source alternatives. The discussion highlights the evolution of language models from Claude Shannon's 1948 model to GPT-3 and beyond, emphasizing the role of RLHF (Reinforcement Lea

x-aiopenaimistral-aianthropic gpt-3gpt-4gpt-5gpt-6
2024-05-22
ALL of AI Engineering in One Place

The upcoming AI Engineer World's Fair in San Francisco from June 25-27 will feature a significantly expanded format with booths, talks, and workshops from top model labs like OpenAI, DeepMind, Anthropic, Mistral, Cohere, HuggingFace, and Character.ai. It includes participation from Microsoft Azure, Amazon AWS, Google Vertex, and major companies such as Nvidia, Salesforce, Mastercard, Palo Alto Networks, and more. The event covers 9 tracks including RAG, multimod

openaigoogle-deepmindanthropicmistral-ai claude-3-sonnetclaude-3
2024-05-16
Cursor reaches >1000 tok/s finetuning Llama3-70b for fast file editing

Cursor, an AI-native IDE, announced a speculative edits algorithm for code editing that surpasses GPT-4 and GPT-4o in accuracy and latency, achieving speeds of over 1000 tokens/s on a 70b model. OpenAI released GPT-4o with multimodal capabilities including audio, vision, and text, noted to be 2x faster and 50% cheaper than GPT-4 turbo, though with mixed coding performance. Anthropic introduced streaming, forced tool use, and vision features for developers.

cursoropenaianthropicgoogle-deepmind gpt-4gpt-4ogpt-4-turbogpt-4o-mini
2024-05-08
OpenAI's PR Campaign?

OpenAI faces user data deletion backlash over its new partnership with StackOverflow amid GDPR complaints and US newspaper lawsuits, while addressing election year concerns with efforts like the Media Manager tool for content opt-in/out by 2025 and source link attribution. Microsoft develops a top-secret airgapped GPT-4 AI service for US intelligence agencies. OpenAI releases the Model Spec outlining responsible AI content generation policies, including NSFW content handling and profanit

openaimicrosoftgoogle-deepmind alphafold-3xlstmgpt-4
Apr 2024 7 issues
2024-04-30
LLMs-as-Juries

OpenAI has rolled out the memory feature to all ChatGPT Plus users and partnered with the Financial Times to license content for AI training. Discussions on OpenAI's profitability arise due to paid training data licensing and potential GPT-4 usage limit reductions. Users report issues with ChatGPT's data cleansing after the memory update. Tutorials and projects include building AI voice assistants and interface agents powered by LLMs. In Stable Diffusion, users seek reali

openaicoherefinancial-times gpt-4gpt-3.5sdxlponyxl
2024-04-24
OpenAI's Instruction Hierarchy for the LLM OS

OpenAI published a paper introducing the concept of privilege levels for LLMs to address prompt injection vulnerabilities, improving defenses by 20-30%. Microsoft released the lightweight Phi-3-mini model with 4K and 128K context lengths. Apple open-sourced the OpenELM language model family with an open training and inference framework. An instruction accuracy benchmark compared 12 models, with Claude 3 Opus, GPT-4 Turbo, and Llama 3 70B performing best. The Rho

openaimicrosoftappledeepseek phi-3-miniopenelmclaude-3-opusgpt-4-turbo
2024-04-22
FineWeb: 15T Tokens, 12 years of CommonCrawl (deduped and filtered, you're welcome)

2024 has seen a significant increase in dataset sizes for training large language models, with Redpajama 2 offering up to 30T tokens, DBRX at 12T tokens, Reka Core/Flash/Edge with 5T tokens, and Llama 3 trained on 15T tokens. Huggingface released an open dataset containing 15T tokens from 12 years of filtered CommonCrawl data, enabling training of models like Llama 3 if compute resources are available. On Reddit, WizardLM-2-8x22b outperform

huggingfacemeta-ai-fairdbrxreka-ai llama-3-70bllama-3wizardlm-2-8x22bclaude-opus
2024-04-18
Meta Llama 3 (8B, 70B)

Meta partially released Llama 3 models including 8B and 70B variants, with a 400B variant still in training, touted as the first GPT-4 level open-source model. Stability AI launched Stable Diffusion 3 API with model weights coming soon, showing competitive realism against Midjourney V6. Boston Dynamics unveiled an electric humanoid robot Atlas, and Microsoft introduced the VASA-1 model generating lifelike talking faces at 40fps on RTX 4090. Mistr

meta-ai-fairstability-aiboston-dynamicsmicrosoft llama-3-8bllama-3-70bllama-3-400bstable-diffusion-3
2024-04-16
Lilian Weng on Video Diffusion

OpenAI expands with a launch in Japan, introduces a Batch API, and partners with Adobe to bring the Sora video model to Premiere Pro. Reka AI releases the Reka Core multimodal language model. WizardLM-2 is released showing impressive performance, and Llama 3 news is anticipated soon. Geoffrey Hinton highlights AI models exhibiting intuition, creativity, and analogy recognition beyond humans. The Devin AI model notably contributes to its own codebase. *

openaiadobereka-ai wizardlm-2llama-3reka-coredevin
2024-04-08
Anime pfp anon eclipses $10k A::B prompting challenge

Victor Taelin issued a $10k challenge to GPT models, initially achieving only 10% success with state-of-the-art models, but community efforts surpassed 90% success within 48 hours, highlighting GPT capabilities and common skill gaps. In Reddit AI communities, Command R Plus (104B) is running quantized on M2 Max hardware via Ollama and llama.cpp forks, with GGUF quantizations released on Huggingface. Streaming text-to-video generation is now available through the *

openaiollamahuggingface command-r-plus-104bstable-diffusion-1.5
2024-04-03
ReALM: Reference Resolution As Language Modeling

Apple is advancing in AI with a new approach called ReALM: Reference Resolution As Language Modeling, which improves understanding of ambiguous references using three contexts and finetunes a smaller FLAN-T5 model that outperforms GPT-4 on this task. In Reddit AI news, an open-source coding agent SWE-agent achieves 12.29% on the SWE-bench benchmark, and RAGFlow introduces a customizable retrieval-augmented generation engine. A new quantization method, QuaRot, enab

appleopenaihugging-facestability-ai flan-t5gpt-4
Mar 2024 5 issues
2024-03-25
Andrew likes Agents

Andrew Ng's The Batch writeup on Agents highlighted the significant improvement in coding benchmark performance when using an iterative agent workflow, with GPT-3.5 wrapped in an agent loop achieving up to 95.1% correctness on HumanEval, surpassing GPT-4 zero-shot at 67.0%. The report also covers new developments in Stable Diffusion models like Cyberrealistic_v40, Platypus XL, and SDXL Lightning for Naruto-style image generation, alongside innovations in LoRA

openaistability-ai gpt-3.5gpt-4cyberrealistic_v40platypus-xl
2024-03-18
Grok-1 in Bio

Grok-1, a 314B parameter Mixture-of-Experts (MoE) model from xAI, has been released under an Apache 2.0 license, sparking discussions on its architecture, finetuning challenges, and performance compared to models like Mixtral and Miqu 70B. Despite its size, its MMLU benchmark performance is currently unimpressive, with expectations that Grok-2 will be more competitive. The model's weights and code are publicly available, encouraging community experimentation. Sam Al

xaimistral-aiperplexity-aigroq grok-1mixtralmiqu-70bclaude-3-opus
2024-03-13
DeepMind SIMA: one AI, 9 games, 600 tasks, vision+language ONLY

DeepMind SIMA is a generalist AI agent for 3D virtual environments evaluated on 600 tasks across 9 games using only screengrabs and natural language instructions, achieving 34% success compared to humans' 60%. The model uses a multimodal Transformer architecture. Andrej Karpathy outlines AI autonomy progression in software engineering, while Arav Srinivas praises Cognition Labs' AI agent demo. François Chollet expresses skepticism about automating software enginee

deepmindcognition-labsdeepgrammodal-labs llama-3claude-3-opusclaude-3gpt-3.5-turbo
2024-03-11
Fixing Gemma

Google's Gemma model was found unstable for finetuning until Daniel Han from Unsloth AI fixed 8 bugs, improving its implementation. Yann LeCun explained technical details of a pseudo-random bit sequence for adaptive equalizers, while François Chollet discussed the low information bandwidth of the human visual system. Arav Srinivas reported that Claude 3 Opus showed no hallucinations in extensive testing, outperforming GPT-4 and Mistral-Large in benchmarks. Reflect

googleunslothanthropicmistral-ai gemmaclaude-3-opusclaude-3mistral-large
2024-03-06
Not much happened today

Anthropic released Claude 3, replacing Claude 2.1 as the default on Perplexity AI, with Claude 3 Opus surpassing GPT-4 in capability. Debate continues on whether Claude 3's performance stems from emergent properties or pattern matching. LangChain and LlamaIndex added support for Claude 3 enabling multimodal and tool-augmented applications. Despite progress, current models still face challenges in out-of-distribution reasoning and robustness. Cohere partnered with Ac

anthropicperplexitylangchainllamaindex claude-3claude-3-opusclaude-3-sonnetgpt-4
Feb 2024 7 issues
2024-02-28
... and welcome AI Twitter!

The AI Twitter discourse from 2/27-28/2024 covers a broad spectrum including ethical considerations highlighted by Margaret Mitchell around Google Gemini's launch, and John Carmack's insights on evolving coding skills in the AI era. Guillaume Lample announced the release of the Mistral Large multilingual model. Discussions also touched on potential leadership changes at Google involving Sundar Pichai, and OpenAI's possible entry into the synthetic data mar

googleopenaiapplestripe mistral-largegoogle-gemini
2024-02-19
Companies liable for AI hallucination is Good Actually for AI Engineers

Air Canada faced a legal ruling requiring it to honor refund policies communicated by its AI chatbot, setting a precedent for corporate liability in AI engineering accuracy. The tribunal ordered a refund of $650.88 CAD plus damages after the chatbot misled a customer about bereavement travel refunds. Meanwhile, AI community discussions highlighted innovations in quantization techniques for GPU inference, Retrieval-Augmented Generation (RAG) and fine-tuning of LLMs, and CUDA o

air-canadahuggingfacemistral-ai mistral-nextlarge-world-modelsorababilong
2024-02-14
AI gets Memory

AI Discords analysis covered 20 guilds, 312 channels, and 6901 messages. The report highlights the divergence of RAG style operations for context and memory, with implementations like MemGPT rolling out in ChatGPT and LangChain. The TheBloke Discord discussed open-source large language models such as the Large World Model with contexts up to 1 million tokens, and the Cohere aya model supporting 101 languages. Roleplay-focused models like Miqu

openailangchaintheblokecohere miqumaid-v2-70bmixtral-8x7b-qloramistral-7bphi-2
2024-02-12
The Dissection of Smaug (72B)

Abacus AI launched Smaug 72B, a large finetune of Qwen 1.0, which remains unchallenged on the Hugging Face Open LLM Leaderboard despite skepticism from Nous Research. LAION introduced a local voice assistant model named Bud-E with a notable demo. The TheBloke Discord community discussed model performance trade-offs between large models like GPT-4 and smaller quantized models, fine-tuning techniques using datasets like WizardLM_evol_instruct_V2_196k and O

abacus-aihugging-facenous-researchlaion smaug-72bqwen-1.0qwen-1.5gpt-4
2024-02-08
Gemini Ultra is out, to mixed reviews

Google released Gemini Ultra as a paid tier for "Gemini Advanced with Ultra 1.0" following the discontinuation of Bard. Reviews noted it is "slightly faster/better than ChatGPT" but with reasoning gaps. The Steam Deck was highlighted as a surprising AI workstation capable of running models like Solar 10.7B. Discussions in AI communities covered topics such as multi-GPU support for OSS Unsloth, training data contamination from OpenAI outputs, ethical concerns over model merging, and n

googleopenaimistral-aihugging-face gemini-ultragemini-advancedsolar-10.7bopenhermes-2.5-mistral-7b
2024-02-05
Less Lazy AI

The AI Discord summaries for early 2024 cover various community discussions and developments. Highlights include 20 guilds, 308 channels, and 10449 messages analyzed, saving an estimated 780 minutes of reading time. Key topics include Polymind Plugin Puzzle integrating PubMed API, roleplay with HamSter v0.2, VRAM challenges in Axolotl training, fine-tuning tips for FLAN-T5, and innovative model merging strategies. The Nous Research AI community discussed G

openaihugging-facenous-researchh2oai hamster-v0.2flan-t5miqu-1-120b-ggufqwen2
2024-02-01
Trust in GPTs at all time low

Discord communities were analyzed with 21 guilds, 312 channels, and 8530 messages reviewed, saving an estimated 628 minutes of reading time. Discussions highlighted challenges with GPTs and the GPT store, including critiques of the knowledge files capability and context management issues. The CUDA MODE Discord was introduced for CUDA coding support. Key conversations in the TheBloke Discord covered Xeon GPU server cost-effectiveness, Llama3 and M

openaihugging-facemistral-ainous-research llama-3mistral-mediumllava-1.6miquella-120b-gguf
Jan 2024 6 issues
2024-01-29
RWKV "Eagle" v5: Your move, Mamba

RWKV v5 Eagle was released with better-than-mistral-7b evaluation results, trading some English performance for multilingual capabilities. The mysterious miqu-1-70b model sparked debate about its origins, possibly a leak or distillation of Mistral Medium or a fine-tuned Llama 2. Discussions highlighted fine-tuning techniques, including the effectiveness of 1,000 high-quality prompts over larger mixed-quality datasets, and tools like Deepspeed, Axolotl, and QLoRA

eleutheraimistral-aihugging-facellamaindex rwkv-v5mistral-7bmiqu-1-70bmistral-medium
2024-01-23
RIP Latent Diffusion, Hello Hourglass Diffusion

Katherine Crowson from Stable Diffusion introduces a hierarchical pure transformer backbone for diffusion-based image generation that efficiently scales to megapixel resolutions with under 600 million parameters, improving upon the original ~900M parameter model. This architecture processes local and global image phenomena separately, enhancing efficiency and resolution without latent steps. Additionally, Meta's Self Rewarding LM paper has inspired lucidrains to begin an implementati

stable-diffusionmeta-ai-fairopenaihugging-face gpt-4latent-diffusion
2024-01-15
1/13-14/2024: Don't sleep on #prompt-engineering

The OpenAI Discord community engaged in diverse discussions including prompt engineering techniques like contrastive Chain of Thought and step back prompting, and explored model merging and mixture-of-experts (MoE) concepts. Philosophical debates on AI consciousness and the ethics of AI-generated voices highlighted concerns about AI sentience and copyright issues. Technical clarifications were made on hyperdimensional vector space models used in modern AI embeddings.

openai
2024-01-10
1/9/2024: Nous Research lands $5m for Open Source AI

Nous Research announced a $5.2 million seed financing focused on Nous-Forge, aiming to embed transformer architecture into chips for powerful servers supporting real-time voice agents and trillion parameter models. Rabbit R1 launched a demo at CES with mixed reactions. OpenAI shipped the GPT store and briefly leaked an upcoming personalization feature. A new paper on Activation Beacon proposes a solution to extend LLMs' context window significantly, with code to b

nous-researchopenairabbit-tech qloraphi-3mixtralollama
2024-01-07
1/6-7/2024: LlaMA Pro - an alternative to PEFT/RAG??

New research papers introduce promising Llama Extensions including TinyLlama, a compact 1.1B parameter model pretrained on about 1 trillion tokens for 3 epochs, and LLaMA Pro, an 8.3B parameter model expanding LLaMA2-7B with additional training on 80 billion tokens of code and math data. LLaMA Pro adds layers to avoid catastrophic forgetting and balances language and code tasks but faces scrutiny for not using newer models like Mistral or Qwen. Meanwhile,

openaimistral-aillamaindexlangchain llama-3llama-3-1-1bllama-3-8-3bgpt-4
2024-01-02
1/2/2024: Smol tweaks to Smol Talk

OpenAI Discord discussions highlight a detailed comparison of AI search engines including Perplexity, Copilot, Bard, and Claude 2, with Bard and Claude 2 trailing behind. Meta AI chatbot by Meta is introduced, available on Instagram and Whatsapp, featuring image generation likened to a free GPT version. Users report multiple browser issues with ChatGPT, including persistent captchas when using VPNs and plugin malfunctions. Debates cover prompt engineering, API usage,

openaimeta-ai-fairperplexity-ai claude-2bardcopilotmeta-ai

2023

Dec 2023 3 issues
2023-12-25
12/25/2023: Nous Hermes 2 Yi 34B for Christmas

Teknium released Nous Hermes 2 on Yi 34B, positioning it as a top open model compared to Mixtral, DeepSeek, and Qwen. Apple introduced Ferret, a new open-source multimodal LLM. Discussions in the Nous Research AI Discord focused on AI model optimization and quantization techniques like AWQ, GPTQ, and AutoAWQ, with insights on proprietary optimization and throughput metrics. Additional highlights include the addition of NucleusX Model to

teknimnous-researchapplemixtral nous-hermes-2yi-34bnucleusxyayi-2
2023-12-18
12/18/2023: Gaslighting Mistral for fun and profit

OpenAI Discord discussions reveal comparisons among language models including GPT-4 Turbo, GPT-3.5 Turbo, Claude 2.1, Claude Instant 1, and Gemini Pro, with GPT-4 Turbo noted for user-centric explanations. Rumors about GPT-4.5 remain unconfirmed, with skepticism prevailing until official announcements. Users discuss technical challenges like slow responses and API issues, and explore role-play prompt techniques to enhance model performance. Ethical concerns about

openaianthropicgoogle-deepmind gpt-4-turbogpt-3.5-turboclaude-2.1claude-instant-1
2023-12-12
12/12/2023: Towards LangChain 0.1

The Langchain rearchitecture has been completed, splitting the repo for better maintainability and scalability, while remaining backwards compatible. Mistral launched a new Discord community, and Anthropic is rumored to be raising another $3 billion. On the OpenAI Discord, discussions covered information leakage in AI training, mixture of experts (MoE) models like mixtral 8x7b, advanced prompt engineering techniques, and issues with ChatGPT performance and

langchainmistral-aianthropicopenai mixtral-8x7bphi-2gpt-3chatgpt