EP223: Ollama vs vLLM vs SGLang
Ollama, vLLM, and SGLang are three primary engines for serving open-weight models, each architected for distinct execution workloads. Ollama relies on a FIFO queue and pre-quantized GGUF models, making it optimal for local prototyping and laptop-scale hardware. vLLM implements continuous batching and PagedAttention to optimize KV cache management and maximize GPU utilization for high-traffic environments. SGLang uses a prefix-aware scheduler and RadixAttention to avoid recomputing prompt prefixes, making it ideal for agent loops, multi-turn chats, and structured outputs. Anthropic plans to watermark AI-generated text to enable reliable identification of model outputs. Rather than standard random sampling, a keyed function uses a secret key and preceding words to constrain candidate token choices during generation. Detection works by verifying whether tokens match the keyed criteria across the text, producing an AI-generation score based on the match rate. The author also notes ongoing issues with false negatives in AI text detection, particularly within technical writing. Agent skills are instruction sets and scripts designed to equip LLM agents with specialized capabilities. As of August 2026, the twelve most-starred skill repositories on GitHub feature a range of augmentations for coding, UI design, planning, and agent communication. These skills include tooling from individual developers as well as organizations like Anthropic, Google, and Multica AI. Together, they assist agents in navigating complex codebases, adopting production-grade practices, and interacting effectively.
閱讀原文 ↗目錄
Ollama vs vLLM vs SGLang
Ollama, vLLM, and SGLang are three primary engines for serving open-weight models, each architected for distinct execution workloads. Ollama relies on a FIFO queue and pre-quantized GGUF models, making it optimal for local prototyping and laptop-scale hardware. vLLM implements continuous batching and PagedAttention to optimize KV cache management and maximize GPU utilization for high-traffic environments. SGLang uses a prefix-aware scheduler and RadixAttention to avoid recomputing prompt prefixes, making it ideal for agent loops, multi-turn chats, and structured outputs.
- Ollama processes requests sequentially through a FIFO queue and runs compressed GGUF models.
- Ollama is targeted at local development, prototyping, and laptop-scale hardware.
- vLLM incorporates continuous batching to add new requests into active batches dynamically.
- vLLM relies on PagedAttention to store the KV cache and is optimized for high-traffic concurrent serving.
- SGLang utilizes a prefix-aware scheduler and RadixAttention to retain and reuse shared prompt prefixes via a radix tree.
- SGLang is tailored for AI agents, multi-turn conversations, and JSON or regex structured outputs.
How does Claude's text watermark work?
Anthropic plans to watermark AI-generated text to enable reliable identification of model outputs. Rather than standard random sampling, a keyed function uses a secret key and preceding words to constrain candidate token choices during generation. Detection works by verifying whether tokens match the keyed criteria across the text, producing an AI-generation score based on the match rate. The author also notes ongoing issues with false negatives in AI text detection, particularly within technical writing.
- Anthropic announced plans to watermark generated text to facilitate AI identification.
- Watermarking alters next-word selection using a keyed function dependent on a secret key and preceding context.
- Tokens where multiple valid word options exist carry the watermark signal.
- Detection evaluates the match rate of words conforming to the keyed selection rule across a piece of text.
- AI text detection techniques currently suffer from false negatives, particularly in technical writing contexts.
Top 12 Agent Skills You Should Know
Agent skills are instruction sets and scripts designed to equip LLM agents with specialized capabilities. As of August 2026, the twelve most-starred skill repositories on GitHub feature a range of augmentations for coding, UI design, planning, and agent communication. These skills include tooling from individual developers as well as organizations like Anthropic, Google, and Multica AI. Together, they assist agents in navigating complex codebases, adopting production-grade practices, and interacting effectively.
- Agent skills consist of instructions and scripts that teach LLM agents specific new capabilities.
- The featured list captures the 12 most-starred agent skill repositories on GitHub as of August 2026.
- Skills range from functional abilities like document creation and codebase knowledge graph generation to behavioral modifications like pre-code planning.
- Major organizations and figures such as Anthropic, Google's Addy Osmani, and Multica AI have developed or contributed to top-starred skills.
Git Workflow: Essential Commands
Git workflows center on tracking and moving code between four distinct locations: the working directory, staging area, local repository, and remote repository. Most workflow issues arise from misunderstanding where code is placed after running specific commands. Core commands facilitate saving work, fetching or checking out repositories, syncing remote changes, and safely stashing uncommitted progress.
- Most practical Git workflows only rely on a small fraction of the available commands.
- Code in Git moves between four main locations: the working directory, staging area, local repository, and remote repository.
- 'git add' stages working directory files, 'git commit' records them to the local repository, and 'git push' uploads them to a remote repository.
- 'git pull' combines the actions of 'git fetch' and 'git merge' into a single step.
- 'git stash' temporarily stores uncommitted changes, which can be restored with 'git stash apply' or restored and deleted with 'git stash pop'.
Apache Kafka vs. RabbitMQ
Apache Kafka and RabbitMQ are messaging technologies built for fundamentally different distributed systems use cases. Kafka operates as a distributed log where messages are appended to partitions, persisted according to retention policies, and pulled by consumers via offsets for replayable, high-throughput event streaming. In contrast, RabbitMQ serves as a message broker where messages are routed through exchanges to queues, pushed to consumers, and deleted upon acknowledgment. Misapplying Kafka as a simple queue or RabbitMQ as an event log undermines their specialized architectures.
- Kafka is designed as a distributed log where producers append messages to partitions based on a retention policy.
- Kafka consumers pull messages using offsets, allowing data to be rewound, replayed, and reprocessed independently.
- RabbitMQ functions as a message broker that routes messages via exchanges into queues using direct, topic, or fanout binding keys.
- RabbitMQ pushes messages to consumers and deletes them once they are acknowledged.
- Using Kafka as a standard task queue or RabbitMQ as an immutable event log is a common architectural anti-pattern.