WebMCP By Google, Clearly Explained!
Superlinked has released the Superlinked Inference Engine (SIE), an open-source tool designed to reduce AI self-hosting costs by approximately 4x. Unlike traditional setups that require separate servers for each model in an agentic pipeline, SIE serves multiple models from a single process on one GPU. It optimizes resource usage by dynamically loading and evicting models based on traffic using a least-recently-used policy. The engine supports over 85 models and is compatible with major vector databases and AI frameworks. WebMCP is a browser API developed by the Chrome and Edge teams that allows websites to explicitly define actions for AI agents. By providing named tools with descriptions and JSON Schema inputs, it eliminates the need for agents to guess functionality from pixels or generic DOM structures. The API operates within the user's existing browser session, ensuring that agents have immediate access to authenticated features without requiring separate API keys. Developers can integrate WebMCP by registering JavaScript objects or by adding specific attributes to existing HTML forms. Sparse random projection serves as an efficient alternative to PCA for high-dimensional datasets, addressing PCA's cubic time complexity which becomes impractical at 1000+ dimensions. The technique projects data into a lower-dimensional space while approximately preserving the pairwise distances between points, a property validated by similar silhouette scores in K-Means clustering experiments. While effective for datasets with over 700-800 features, the method can be unstable for low-dimensional data and requires tuning the target dimension as a hyperparameter. This principle is also applied in LLM fine-tuning through the VeRA technique, which uses frozen, shared random matrices to reduce trainable parameters.
閱讀原文 ↗目錄
Researchers built a new AI inference engine
Superlinked has released the Superlinked Inference Engine (SIE), an open-source tool designed to reduce AI self-hosting costs by approximately 4x. Unlike traditional setups that require separate servers for each model in an agentic pipeline, SIE serves multiple models from a single process on one GPU. It optimizes resource usage by dynamically loading and evicting models based on traffic using a least-recently-used policy. The engine supports over 85 models and is compatible with major vector databases and AI frameworks.
- SIE reduces self-hosting costs by consolidating multiple models onto a single GPU process.
- The engine supports over 20 model architectures and 85+ specific models, including embedders, rerankers, and LLMs.
- It uses a least-recently-used (LRU) eviction strategy to manage GPU memory dynamically based on traffic.
- SIE is a drop-in replacement for the OpenAI API and is licensed under Apache 2.0.
- The tool integrates with vector databases like Qdrant, Weaviate, Chroma, and LanceDB, as well as LangChain and LlamaIndex.
WebMCP by Google, clearly explained!
WebMCP is a browser API developed by the Chrome and Edge teams that allows websites to explicitly define actions for AI agents. By providing named tools with descriptions and JSON Schema inputs, it eliminates the need for agents to guess functionality from pixels or generic DOM structures. The API operates within the user's existing browser session, ensuring that agents have immediate access to authenticated features without requiring separate API keys. Developers can integrate WebMCP by registering JavaScript objects or by adding specific attributes to existing HTML forms.
- WebMCP enables websites to expose specific functionalities as machine-readable tools with plain-English descriptions and typed inputs.
- The API uses JSON Schema for input validation, ensuring compatibility with major models like Claude, GPT, and Gemini.
- Actions run within the user's active browser tab, leveraging existing login sessions and removing the need for separate authentication tokens.
- Available tools can dynamically update based on the user's state, such as showing more actions after a user signs in.
- Implementation can be done via a JavaScript registerTool method or by adding toolname and tooldescription attributes to HTML forms.
- WebMCP is positioned as a more reliable and token-efficient alternative to pixel-based 'computer use' or generic browser automation.
Sparse random projections
Sparse random projection serves as an efficient alternative to PCA for high-dimensional datasets, addressing PCA's cubic time complexity which becomes impractical at 1000+ dimensions. The technique projects data into a lower-dimensional space while approximately preserving the pairwise distances between points, a property validated by similar silhouette scores in K-Means clustering experiments. While effective for datasets with over 700-800 features, the method can be unstable for low-dimensional data and requires tuning the target dimension as a hyperparameter. This principle is also applied in LLM fine-tuning through the VeRA technique, which uses frozen, shared random matrices to reduce trainable parameters.
- PCA's cubic relationship with dimensions makes it unsuitable for very high-dimensional datasets.
- Sparse random projection preserves the distance between data points while reducing dimensionality.
- K-Means clustering performance, measured by silhouette score, remains stable when using random projections until the target dimension becomes too low.
- The target dimension for projection is a hyperparameter that must be tuned to balance efficiency and clustering quality.
- Random projections are recommended for datasets with 700-800+ features but are unstable for data with fewer than 100 dimensions.
- VeRA is an LLM fine-tuning method that utilizes frozen, shared random matrices to minimize trainable parameters compared to LoRA.