RAG & Fine-tuning in LLMs
MongoDB University has launched a free AI Skill Badges program designed to help developers build production-grade AI applications. The curriculum focuses on hands-on learning and the creation of functional systems rather than theoretical concepts. Key tracks include auto-embedding AI models, agentic memory design, and semantic search optimization. These badges cover essential technologies including LangGraph, Voyage AI, and MongoDB's native vector search capabilities. RAG and fine-tuning are distinct but complementary techniques for optimizing Large Language Models (LLMs) in production. RAG provides external context at inference time without modifying model weights, making it ideal for dynamic or document-specific knowledge. Fine-tuning updates model weights offline to adapt the model's behavior, style, or specialized reasoning patterns. While often viewed as alternatives, they are frequently combined to create robust systems that require both factual accuracy and specific behavioral traits. This section demonstrates how to build a local browser automation system using an open-source AI stack. The architecture leverages CrewAI for multi-agent orchestration and Stagehand for autonomous web navigation and interaction. The workflow is divided among specialized agents for planning, execution, and response synthesis, all coordinated through CrewAI Flows.
閱讀原文 ↗目錄
Free AI Skill Badges to learn production AI capabilities
MongoDB University has launched a free AI Skill Badges program designed to help developers build production-grade AI applications. The curriculum focuses on hands-on learning and the creation of functional systems rather than theoretical concepts. Key tracks include auto-embedding AI models, agentic memory design, and semantic search optimization. These badges cover essential technologies including LangGraph, Voyage AI, and MongoDB's native vector search capabilities.
- MongoDB University offers free AI Skill Badges focused on practical, production-grade AI implementation.
- The 'Auto-embedding AI Models' track teaches how to generate vector embeddings directly within MongoDB.
- The 'Agentic Memory' badge utilizes LangGraph and Voyage AI to build short and long-term memory for AI agents.
- Voyage AI is used to build two-step retrieval pipelines for semantic search and RAG to optimize latency and cost.
- The program includes foundational tracks such as Vector Search Fundamentals and AI Data Strategy.
- Every badge focuses on building working systems rather than abstract walkthroughs.
RAG & Fine-tuning, explained visually
RAG and fine-tuning are distinct but complementary techniques for optimizing Large Language Models (LLMs) in production. RAG provides external context at inference time without modifying model weights, making it ideal for dynamic or document-specific knowledge. Fine-tuning updates model weights offline to adapt the model's behavior, style, or specialized reasoning patterns. While often viewed as alternatives, they are frequently combined to create robust systems that require both factual accuracy and specific behavioral traits.
- RAG operates at inference time by providing a 'cheat sheet' of context to the LLM without changing its weights.
- Fine-tuning happens offline and updates model weights to change default behavior, tone, and vocabulary.
- RAG is the preferred solution for tasks requiring access to frequently updated external documents or product catalogs.
- Fine-tuning is necessary for adopting internal jargon, specific writing styles, or domain-specific reasoning patterns.
- Production systems often use both RAG and fine-tuning together to balance knowledge retrieval and brand voice.
- Advanced retrieval methods include ColBERT for multivector retrieval and ColPali for complex document processing.
- Parameter-efficient fine-tuning techniques like LoRA and DoRA allow for behavioral adaptation with reduced overhead.
Building a Browser Automation Agent
This section demonstrates how to build a local browser automation system using an open-source AI stack. The architecture leverages CrewAI for multi-agent orchestration and Stagehand for autonomous web navigation and interaction. The workflow is divided among specialized agents for planning, execution, and response synthesis, all coordinated through CrewAI Flows.
- Stagehand is an open-source tool providing computer-use agentic capabilities for browser automation.
- CrewAI serves as the orchestration layer to manage agent workflows and state transitions.
- Ollama is utilized to run the gpt-oss model locally for the automation tasks.
- The system architecture uses a three-agent structure: a Planner Agent, a Browser Automation Agent, and a Response Agent.
- CrewAI Flows is the specific feature used to connect agents into a functional workflow.
- The Browser Automation Agent can autonomously navigate URLs, perform actions, and extract data from web pages.
ANN search using inverted file index
Approximate nearest neighbor (ANN) search algorithms are designed to overcome the inefficiency of exhaustive kNN searches on large datasets. The Inverted File Index (IFV) is a specific ANN technique that partitions data into clusters using k-means and maps points to their respective centroids. During a search, the algorithm only queries data points within the partition of the closest centroid, significantly reducing time complexity. While this method drastically improves latency, it introduces a trade-off where some nearest neighbors may be missed if they reside in different partitions.
- kNN search has a time complexity of O(ND), making it inefficient for large-scale data.
- Inverted File Index (IFV) reduces search space by partitioning data into K clusters using k-means.
- The search complexity of IFV is O(KD + ND/K), which can be significantly faster than exhaustive search.
- In a scenario with 10 million data points and 100 partitions, IFV can be nearly 100 times faster than kNN.
- ANN techniques prioritize reduced latency over perfect accuracy, accepting the risk of missing some close data points.