Turn Any Website Into a Custom API in Claude Code
Bright Data has introduced Scraper Studio to its CLI, enabling the creation of custom web scrapers through natural language prompts. This tool addresses the limitations of standard utilities like curl or Claude Code's built-in web tools, which often fail due to bot detection or lack of JavaScript support. Scraper Studio automatically generates callable APIs and collectors, providing a scalable way to extract structured data from any website. It also includes an auto-repair feature to maintain scraper functionality when site layouts change. The Hermes Agent masterclass is a 48-minute video guide providing instructions on customizing and understanding the Hermes Agent. It covers technical components including self-evolving skills, a three-tier memory architecture, and GEPA optimization. The tutorial also explains how to scale the deployment from a single agent to a group of ten agents working autonomously. This section explains three primary architectural approaches for pairwise sentence scoring in NLP: Cross-encoders, Bi-encoders, and ColBERT. Cross-encoders provide high semantic expressiveness by processing query-document pairs together but lack scalability for large datasets. Bi-encoders offer high scalability through offline document encoding but lose token-level interactions, while ColBERT bridges these methods using a late interaction mechanism that balances performance and efficiency.
閱讀原文 ↗目錄
Turn any website into a custom API
Bright Data has introduced Scraper Studio to its CLI, enabling the creation of custom web scrapers through natural language prompts. This tool addresses the limitations of standard utilities like curl or Claude Code's built-in web tools, which often fail due to bot detection or lack of JavaScript support. Scraper Studio automatically generates callable APIs and collectors, providing a scalable way to extract structured data from any website. It also includes an auto-repair feature to maintain scraper functionality when site layouts change.
- Claude Code's built-in web_fetch tool is restricted to 125-character summaries and cannot retrieve full page content.
- Bright Data CLI manages complex scraping tasks including CAPTCHA solving, bot detection, and browser rendering.
- Scraper Studio allows users to define data fields in natural language to generate custom scrapers and APIs.
- The tool includes an auto-repair mechanism that updates scrapers automatically when a website's structure changes.
- Bright Data provides pre-built extractors for over 40 major platforms such as Amazon, LinkedIn, and TikTok.
- The bdata CLI can be used within Claude Code to automate the generation of Python scripts for data collection.
Hermes agent masterclass
The Hermes Agent masterclass is a 48-minute video guide providing instructions on customizing and understanding the Hermes Agent. It covers technical components including self-evolving skills, a three-tier memory architecture, and GEPA optimization. The tutorial also explains how to scale the deployment from a single agent to a group of ten agents working autonomously.
- The masterclass is a 48-minute video guide for Hermes Agent.
- Hermes Agent utilizes a three-tier memory system.
- The system incorporates GEPA optimization.
- Hermes Agent supports self-evolving skills.
- The guide explains how to scale from 1 to 10 agents working 24/7.
Visual guide to Bi-encoders, Cross-encoders & ColBERT
This section explains three primary architectural approaches for pairwise sentence scoring in NLP: Cross-encoders, Bi-encoders, and ColBERT. Cross-encoders provide high semantic expressiveness by processing query-document pairs together but lack scalability for large datasets. Bi-encoders offer high scalability through offline document encoding but lose token-level interactions, while ColBERT bridges these methods using a late interaction mechanism that balances performance and efficiency.
- Cross-encoders concatenate query and document text into a single BERT-like encoder for high accuracy but require a forward pass for every document in a collection.
- Bi-encoders encode queries and documents separately, allowing for pre-computed document embeddings and fast cosine similarity searches.
- ColBERT utilizes a late interaction mechanism that computes similarity scores between all query and document tokens to maintain semantic expressiveness.
- ColBERT achieves scalability similar to bi-encoders because document embeddings can still be computed offline.
- These scoring methods are foundational components for RAG systems, QA systems, and duplicate text detection.
- ColBERTv2 is an improved version of the original ColBERT architecture designed to enhance RAG systems.