← 回到 Reading
ByteByteGo 2026-08-10

How to Fight Clickbait: Meta, LinkedIn & YouTube Case Studies

Large feed recommendation systems rely on a two-step pipeline composed of retrieval and ranking. Retrieval historically relied on behavioral signals such as clicks and reactions, scaling well across hundreds of millions of candidates. However, optimizing directly for interaction counts favors engagement bait because engagement is an imperfect proxy for true relevance. A durable fix requires altering retrieval itself, transitioning from tracking behavior to evaluating meaning. Semantic retrieval offers an alternative to interaction-based retrieval by finding content through its underlying meaning rather than keyword overlap. This approach relies on embeddings that position related items near each other in a high-dimensional space based on relationships learned during model training. Platforms implement this via a dual-encoder or two-tower design, allowing post embeddings to be precomputed while user embeddings are generated dynamically at request time. A nearest-neighbor search quickly matches users with relevant posts across massive datasets. Major platforms like LinkedIn, Meta, and YouTube use this architecture across search, product recommendations, and retrieval-augmented generation. LinkedIn replaced five separate feed retrieval pipelines with a single unified retrieval model fine-tuned from Meta's LLaMA-3. The system operates as a dual encoder, mapping members and content into a shared embedding space to support nearest-neighbor retrieval with sub-50-millisecond latency. To feed tabular signals into a language model, LinkedIn implemented a prompt library that serializes structured metadata into formatted text sequences. The team also found that transforming raw numeric metrics into ranked percentile buckets improved retrieval accuracy by roughly fifteen percent over using raw counts.

閱讀原文 ↗
目錄 8 段
  1. 01Engagement Signals
  2. 02Semantic Retrieval
  3. 03Unified Retrieval
  4. 04Ranking Funnels
  5. 05Generative Retrieval
  6. 06Cold Start
  7. 07Design Tradeoffs
  8. 08Conclusion

Engagement Signals

Large feed recommendation systems rely on a two-step pipeline composed of retrieval and ranking. Retrieval historically relied on behavioral signals such as clicks and reactions, scaling well across hundreds of millions of candidates. However, optimizing directly for interaction counts favors engagement bait because engagement is an imperfect proxy for true relevance. A durable fix requires altering retrieval itself, transitioning from tracking behavior to evaluating meaning.

  • Large feed architectures depend on a two-step pipeline of candidate retrieval followed by ranking.
  • The retrieval stage reduces hundreds of millions of candidates to roughly one thousand using low-cost operations.
  • Ranking allocates significant compute to order the surviving candidates into the final user feed.
  • Optimizing retrieval strictly for interaction counts systematically elevates engagement bait.
  • Demotions and heuristic rules merely treat the symptoms of engagement bait without addressing retrieval mechanics.
  • A durable architectural improvement shifts retrieval objectives from user behavior to content meaning.

Semantic Retrieval

Semantic retrieval offers an alternative to interaction-based retrieval by finding content through its underlying meaning rather than keyword overlap. This approach relies on embeddings that position related items near each other in a high-dimensional space based on relationships learned during model training. Platforms implement this via a dual-encoder or two-tower design, allowing post embeddings to be precomputed while user embeddings are generated dynamically at request time. A nearest-neighbor search quickly matches users with relevant posts across massive datasets. Major platforms like LinkedIn, Meta, and YouTube use this architecture across search, product recommendations, and retrieval-augmented generation.

  • Embeddings represent content as high-dimensional vectors where semantically related items are positioned close to one another even without shared keywords.
  • A two-tower or dual-encoder architecture uses separate encoders for the user profile/activity and the content items.
  • Item embeddings can be precomputed and indexed offline, requiring only the user embedding and nearest-neighbor search at request time for fast retrieval.
  • Semantic retrieval powered by language models matches concepts based on training associations rather than surface-level token matching.
  • Two-tower retrieval patterns extend beyond social feeds to search ranking, retrieval-augmented generation (RAG), and product recommendations.
  • LinkedIn, Meta, and YouTube have each implemented semantic retrieval systems with divergent architectural solutions.

Unified Retrieval

LinkedIn replaced five separate feed retrieval pipelines with a single unified retrieval model fine-tuned from Meta's LLaMA-3. The system operates as a dual encoder, mapping members and content into a shared embedding space to support nearest-neighbor retrieval with sub-50-millisecond latency. To feed tabular signals into a language model, LinkedIn implemented a prompt library that serializes structured metadata into formatted text sequences. The team also found that transforming raw numeric metrics into ranked percentile buckets improved retrieval accuracy by roughly fifteen percent over using raw counts.

  • LinkedIn consolidated five independent retrieval pipelines into a single model fine-tuned on LLaMA-3 in March 2026.
  • The fine-tuned model functions as a dual encoder that conducts nearest-neighbor retrieval at sub-50-millisecond latency.
  • A prompt library converts structured member and post attributes into templated text representations for the language model.
  • Passing raw integer counts directly into prompt contexts yielded poor correlation with relevance scores due to arbitrary tokenization.
  • Transforming raw engagement counts into ranked percentile buckets increased retrieval accuracy by approximately 15 percent.

Ranking Funnels

Meta designs Instagram recommendation pipeline as a multi-stage funnel supported by an ecosystem of over a thousand models. The architecture moves candidate items through retrieval, early-stage ranking with a two-tower model, late-stage multi-action prediction, a value model, and final adjustments for diversity and integrity. While this specialization simplifies tuning competing objectives like integrity and fairness, it increases operational complexity relative to consolidated approaches like LinkedIn.

  • Meta operates Instagram recommendation system using an ecosystem of more than 1,000 models arranged in a multi-stage funnel.
  • The funnel consists of four distinct steps: candidate retrieval, early-stage ranking via a lightweight two-tower model, late-stage ranking via a heavier model, and a final pass for diversity and integrity.
  • The late-stage ranking model predicts multiple user actions simultaneously, which are subsequently merged into a single composite score by a value model.
  • The scoring system incorporates negative user signals, such as selecting 'See Fewer Posts Like This', alongside positive signals like saves.
  • Multi-stage funnels simplify tuning and auditing disparate objectives like fairness and integrity, but they carry high operational complexity.

Generative Retrieval

YouTube developed a generative retrieval system called PLUM that completely eliminates traditional search indexes from the retrieval process. PLUM assigns each video a content-derived Semantic ID and adapts a pretrained Gemini family language model to generate these IDs based on user history using beam search. While this introduces a hallucination risk of generating invalid IDs, fine-tuning keeps that failure rate below five percent. The system notably improves coverage of long-tail content, boosts YouTube Shorts click-through rates by 4.96 percent, and shifts parameter storage from massive embedding tables into the neural network itself.

  • PLUM replaces traditional search indexes by treating retrieval as a sequence generation task.
  • Videos are represented by Semantic IDs, which are short sequences of discrete codes derived from video content rather than random identifiers.
  • PLUM adapts a pretrained Gemini family language model by adding Semantic IDs to its vocabulary and fine-tuning on video metadata and user activity.
  • The model decodes candidate Semantic IDs using beam search given a user's recent history, mapping generated codes back to real videos.
  • A failure mode of generating nonexistent IDs is kept under five percent after fine-tuning.
  • PLUM improved YouTube Shorts panel click-through by 4.96 percent and significantly enhanced coverage of long-tail videos.
  • The system shifts the parameter footprint from large embedding tables directly into the neural network.

Cold Start

Recommendation systems encounter the cold-start problem when new users lack sufficient interaction history for behavioral retrieval. Semantic retrieval overcomes this by utilizing language model pretraining associations to infer interests directly from profile text. However, relying on these associations can also result in incorrect or stereotypical assumptions when profiles are sparse. LinkedIn's deployment confirmed this dynamic, showing modest overall lift but concentrated gains among new and low-connection users.

  • Behavioral retrieval struggles with cold-start users because it depends on historical interaction signals.
  • Semantic retrieval uses language model pretraining associations to infer user interests from profile metadata before any user clicks occur.
  • A key drawback of semantic retrieval is the risk of inaccurate inferences or stereotyping derived from sparse profile data.
  • LinkedIn's implementation yielded modest overall lift, with improvements concentrated specifically among new and low-connection members.

Design Tradeoffs

System design in recommendation and retrieval involves fundamental tradeoffs between consolidated single models and specialized multi-model funnels. While single-model architectures simplify maintenance and align retrieval with ranking, multi-model architectures provide independent objective control and modular rollback capabilities. Additionally, upgrading to language-model embeddings increases computational demands across data pipelines and serving paths rather than within the model itself. Finally, teams must weigh generative retrieval's compactness against its risk of generating non-existent IDs, whereas index-based retrieval prevents hallucinated IDs at the expense of index maintenance.

  • Single-model architectures simplify maintenance and align retrieval with ranking, whereas multi-model funnels enable granular control and isolated rollbacks across specialized components.
  • Language-model embeddings provide richer representations than lightweight alternatives but require significantly more compute to generate and serve.
  • System bottlenecks for embedding workflows frequently shift away from the model toward data pipelines, feature representation, and serving paths.
  • Generative retrieval eliminates large embedding tables by using compact item codes, but risks hallucinating non-existent item identifiers.
  • Index-based retrieval avoids hallucinated item identifiers but demands continuous storage and maintenance for the index.
  • Semantic retrieval diminishes the effectiveness of bait content without completely preventing content engineered around high-value topics.

Conclusion

Social media feeds are shifting retrieval architectures from behavioral engagement optimization toward semantic meaning, diminishing the efficacy of engagement bait. Although LinkedIn, Meta, and YouTube share this underlying direction, each implemented distinct architectural solutions suited to their data structures. LinkedIn unified five retrieval systems into a single dual-encoder language model over a shared embedding space, Meta retained a staged funnel of specialized models with multi-objective optimization, and YouTube bypassed retrieval indexing by generating item identifiers directly with an adapted language model. Ultimately, architectural design diverges because text-rich networks, multi-objective media platforms, and massive video corpora present vastly different optimization constraints.

  • Retrieval architectures are pivoting from behavioral engagement metrics toward semantic meaning to curb engagement bait.
  • LinkedIn consolidated five separate retrieval systems into a single language-model dual encoder over a unified embedding space.
  • Meta uses a staged funnel powered by a large family of specialized models combined with a multi-objective value model.
  • YouTube generates next-item identifiers directly via an adapted language model, eliminating traditional retrieval indexing.
  • Architectural variations across platforms are primarily driven by differences in data modality, corpus size, and platform objectives.