- Ranking Model Iterations
- Feature Engineering
- Recommendation Pipelines
- Causal Inference and Debiasing
- LLMs for Recommendation
- Large-scale + Real-time + Deep Learning
- Reference
Ranking Model Iterations
The main challenges of modern recommender systems lie in multi-domain, multi-format, and multi-objective fusion.
- Multi-domain and Multi-format
- Multi-scenario coverage: Live streaming, e-commerce lists, content feeds.
- Content diversity: Photo carousels, text-image mixes, horizontal and vertical video formats.
- Diverse interactive feedbacks: Supporting deep reading inside posts, instant in-feed interactions (accessing profile, shares, comments, social behaviors).
- Multi-objective Balancing: Aligned with Product Positioning
- TikTok: Efficiency-first logic.
- Ecosystem balance: promoting viral hits while preventing filter bubbles.
- Balancing commercial value and user value.
- Kuaishou:
- Public-private traffic balance (preserving the creator-fan community while breaking out to new users).
- Lower-tier market support (incubating local/niche merchants).
- Xiaohongshu:
- Community culture: UGC preservation, maintaining user trust.
- “Useful” mindshare: discoverability of high-quality content, knowledge precipitation.
- Weibo:
- Hot event response: stability under burst traffic.
- Social graph cultivation: recommendation as a bridge to social relationships.
- E-commerce (Taobao/Pinduoduo):
- Balancing immediate transactions and Customer Lifetime Value (LTV).
- Merging social loops (“Explore” tab) with transaction values.
- TikTok: Efficiency-first logic.
Multi-objective Fusion
- Evolution of Fusion Methods
- Early Phase: Static formula blending + offline grid search.
- Mid-term Phase: Reinforcement learning for dynamic parameter search.
- Segmenting online traffic into small pools.
- Generating new parameters based on active performance.
- Gathering user feedback to update parameters.
- User preferences shift quickly (e.g., weekends vs. weekdays).
- Parameter fusion must reflect user preferences in real-time.
- Core lies in reward calculations, using Covariance Matrix Adaptation (CMA-ES) or Evolutionary Strategies (ES).
- Optimization Tips:
- User usage has periodic patterns; reset/initialize daily.
- Perform prior analysis before applying differentiated fusion parameters.
- Inject anomaly detection to ensure stable parameter updates.
- Late Phase: Blending formula optimization.
- Latest Trend: outputting a fusion score via models.
- Upgrading to multi-task learning, performing objective fusion using model layers.
- Model-based fusion captures complex non-linear relationships, yielding superior representation power.
- Achieves personalized fusion—the calculated weights are unique to each user.
- Latest Trend: outputting a fusion score via models.
Multi-Task Learning
- Challenges in Multi-task Modeling
- Loss conflicts and the “seesaw effect” across multiple objectives.
- Inconsistent sample spaces (e.g., clicks vs. conversions).
- Balancing task-specific losses.
- Common Multi-task Architectures
- MMoE (Multi-gate Mixture-of-Experts)
- Core: Sharing bottom expert networks, and designing separate gating networks for each task.
- Pros: Gating networks learn relationships across tasks, balancing sharing and specialization.
- Application: Widely used in CTR, CVR multi-objective prediction.
- SNR (Sub-Network Routing)
- Core: Dynamic routing mechanisms, selecting different expert sub-networks per sample.
- Pros: More flexible than MMoE, reducing negative transfer.
- Innovation: Routing mechanism optimizes expert utilization.
- DMT-GRU (Deep Multitask GRU)
- Core: Combining GRU structures with multi-task frameworks.
- Pros: Better captures sequential information and task-specific temporal dependencies.
- Application: Suitable for sequence-dependent multi-task scenarios.
- MM (Mixture of Multitask)
- Core: Adding fusion layers on top of SNR.
- Pros: Better task-specific knowledge transfer; stronger expressive power.
- Features: High capacity, integrating multiple multi-task learning techniques.
- Evolution of Multi-task Modeling:
- Technical Roadmap: Hard parameter sharing MMoE SNR/PLE.
- Team Practice: Adopting SNR models with two key optimizations:
- Simplifying transformation structures within experts.
- Combining shared experts and task-exclusive experts.
- Expert Configuration: Designing experts based on business feedback and estimation bias analyses.
- MMoE (Multi-gate Mixture-of-Experts)
- Production Experience
- Methods like PCGrad and UWL show gains on test datasets but suffer from performance decay in production over time.
- Empirical parameter tuning in online learning environments is often more stable.
- Implementing MMoE on its own can bring significant business gains.
Multi-Domain Modeling
Multi-domain modeling solves data distribution shifts across different scenarios. Compared to multi-task learning, the motivations differ:
- Multi-task Learning: Primarily addresses target sparsity (e.g., rare conversions).
- Multi-domain Modeling: Resolves knowledge transfer from large scenarios to small scenarios.
Key challenges across multiple recommendation scenarios:
- Huge scale differences: Small scenarios struggle to converge due to sparse training data.
- Even in scenarios of similar scale, knowledge transfer can bring business gains.
- Balancing shared knowledge with domain-specific characteristics.
Multi-domain modeling is a recent research hotspot, sharing many technical implementations with multi-task learning.
Multi-domain models typically add a Slot-gate layer on top of multi-task architectures. This layer enables the same embedding to express different actions depending on the domain. The output of the Slot-gate routes in three directions:
- Connecting expert networks.
- Connecting target tasks.
- Connecting features.
In practice, the main model employs SNR (Sub-Network Routing) to replace CGC (Customized Gate Control), aligning with multi-task modeling iterations.
Scenarios:
- Homepage recommendation: Trending Feed.
- Discover tab: Hot Topic Feed.
The overall architecture resembles SNR, designing three target towers at the top: Click, Interaction, and Stay Time. These three target towers are further split into six concrete objectives across the two scenarios.
Furthermore, we added an Embedding Transform Layer. Its difference from the Slot-gate is:
- Slot-gate: Identifies feature importance differences across scenarios.
- Embedding Transform Layer: Resolves representation mismatches in embedding spaces across domains, performing embedding mappings.
This design is highly effective for processing features with mismatched dimensions across scenarios, promoting cross-domain knowledge transfer.
Feature Engineering
Interest Representation and Behavior Sequence Modeling
- Key Technical Evolution
- DIN (Deep Interest Network)
- Core: Building multiple sequences for different behaviors, introducing attention mechanisms.
- Features: Using local activation units to learn interest distributions relative to the candidate item.
- Results: Implemented in trending fine ranking pipelines, bringing significant gains.
- SIM (Search Interest Model)
- Iteration on top of DIN.
- More suitable for interest modeling under search scenarios.
- DMT (Deep Multifaceted Transformers)
- Core: Applying Transformer architectures to multi-task learning.
- Practice: Simplified the DMT model, removing bias modules and substituting MMoE with SNR.
- Results: Achieved business gains upon rollout.
- DIN (Deep Interest Network)
- Sequence Modeling Optimization Directions
- Multi-sequence Fusion (Multi-DIN)
- Method: Expanding multiple behavior sequences, using candidate item features (mid/tag/authorid) as the query.
- Pipeline: Separate attention per sequence extract interest representations concatenate other features feed into multi-task ranking.
- Long-sequence Modeling
- Experiment: Expanding click/duration/interaction sequences from 20 to 50 items.
- Conclusion: Yields superior results, aligning with academic literature.
- Tradeoff: Requires higher computing costs.
- Lifecycle-level Ultra-long Sequence Modeling
- Difference from standard long sequences:
- Constructing user long behavior sequence features offline.
- Finding corresponding features via search indexing to generate embeddings.
- Modeling the main network and ultra-long sequence network separately.
- Business Value Evaluation:
- Limited value in fast-paced platforms like Weibo (where user interest shifts rapidly).
- More valuable for low-frequency or returned users.
- Difference from standard long sequences:
- Multi-sequence Fusion (Multi-DIN)
- Technical Selection Logic
- Sequence length must align with business pace: fast-paced feeds should not use overly long sequences.
- Balance computing cost and performance gains: longer sequences increase serving overhead.
- Differentiated handling: treat active power users and returned users differently.
Feature Construction
Challenges and practical experience in feature engineering within large-scale models:
- Adding features that look good in theory does not always yield expected online gains.
- Large-scale models already contain massive ID features, which capture user preferences well.
- Basic statistical features face diminishing marginal returns under these conditions.
Common Feature Engineering Approaches:
- Matching Features: Detailed cross-statistics between users and items, content categories, and creators. Highly effective.
- Multimodal Features: Resolves sparse user behavior logs for low-frequency/long-tail items.
- Method 1: Direct fusion of multimodal embeddings.
- Freezing bottom embedding gradients, only updating the top MLP.
- Pros: Retains full semantic information.
- Cons: Increases model complexity; requires spatial transforms and feature importance matching.
- Method 2: Multimodal clustering ID-ization.
- Clustering multimodal features offline and feeding clustering IDs into the model.
- Pros: Low model complexity; simple online serving; captures ~90% of the value.
- Combined with statistical features of clustering IDs to boost performance.
- Cons: Losses granular semantic details.
- Method 1: Direct fusion of multimodal embeddings.
- Feature Crossing: Co-action Methods
- Motivation: Traditional cross methods (DeepFM, Wide&Deep) show subpar gains.
- Analysis: Shared embedding spaces between cross-features and DNN layers cause updates to conflict.
- Solution: Allocating independent storage spaces for cross-features.
- Results: Expanded representation space, yielding online gains.
Recommendation Pipelines
Pipeline Consistency
Consistency between pre-ranking and fine ranking is a critical bottleneck:
- Pre-ranking Truncation Issues
- Pre-ranking trims candidates from thousands to ~1000.
- If pre-ranking and fine ranking represent user interests differently, high-scoring candidates in fine ranking might be filtered out early.
- Improving consistency directly drives overall business metrics.
- Sources of Discrepancy
- Feature systems differ: pre-ranking features are highly simplified.
- Model structures differ: pre-ranking favors lightweight structures.
- Feature interaction timing: pre-ranking crosses features late, limiting representation power.
- Technical Evolution Roads
- Dual-tower (Two-tower) Line:
- Pros: High computing efficiency; suitable for large candidate pools.
- Cons: Late feature interactions limit representation power.
- Attempts: DSSM-autowide and other cross-tower structures, similar to DeepFM.
- DNN Line (Post-2022):
- Pros: Stronger representation power; higher scoring quality.
- Cons: Heavy load on engineering; requires feature pruning, network trimming, and caching.
- Tradeoff: Although the number of processed candidates per query decreases, it is accepted due to higher scoring quality.
- Dual-tower (Two-tower) Line:
- Team Experience Summary
- Enhancing dual-tower models (e.g., late feature crossing) brings limited gains.
- Multi-task pre-ranking is still bottlenecked by the dual-tower structure.
- Migrating to DNN architectures is the key path to boosting pre-ranking accuracy.
- Enhancing pre-ranking and fine ranking consistency is a vital optimization direction.
This consistency challenge essentially reflects information loss across multi-stage recommendation funnels, requiring an optimal balance between representation accuracy and computing budgets.
Cascade Models
Cascade models are effective solutions to pre-ranking/fine ranking consistency:
- Architectural Advantages
- Employing DNN and cascade Stacking models to achieve internal “coarse-to-fine” two-stage pre-ranking.
- Using dual towers for fast filtering, followed by a lightweight DNN for final scoring.
- The DNN accommodates complex structures, adapting faster to user interest shifts.
- Handling Large Candidate Pools
- Plays a vital role in recommendation frameworks, filtering from large pools efficiently.
- Bypasses the computing bottleneck of applying full DNNs directly on massive candidates.
- Sample Construction Strategies
- Core lies in designing appropriate training samples.
- Sampling across recommendation funnel stages (millions catalog thousands retrieved hundreds ranked dozens exposed single-digit clicks).
- Combining sample pairs of varying difficulty to boost model learning.
Cascade models combined with global negative sampling bring significant gains, balancing representation power and computing costs.
Causal Inference and Debiasing
Causal Inference Value:
- Balancing Personalization
- Solving conflicts between popular item biases and niche individual interests.
- Boosting the platform’s personalization accuracy.
- Preventing models from over-recommending popular items that clash with the user’s active interests.
- Implementation Methods
- Constructing specialized pairwise samples: clicked low-popularity items vs. unclicked high-popularity items.
- Designing Bayesian loss functions to guide model training.
- Utilizing counterfactual reasoning to isolate true interest causal paths.
- Application Strategies
- Applying causal inference in retrieval and pre-ranking yields superior gains compared to fine ranking.
- Brings significant improvements to model components with weaker personalization power.
- Displays an inverse relationship with model complexity and representation capacity.
Importance of Debiasing:
- Data Bias Mismatch
- Recommender systems are prone to exposure, position, and popularity biases.
- These biases cause models to learn from skewed behavior logs rather than true user interests.
- Left unchecked, it creates a feedback loop (Mathew effect), degrading the long-tail content ecosystem.
- Core Value of Debiasing
- Breaking positive feedback loops, preventing filter bubbles.
- Enhancing content diversity, opening up new interests.
- Providing fair exposure opportunities for creators, fostering a healthy ecosystem.
- Implementation Challenges
- Balancing debiasing with short-term recommendation precision.
- Combining causal graphs and intervention models to isolate and eliminate harmful biases.
- Utilizing counterfactual learning and Propensity Score Matching (PSM) to achieve effective debiasing.
Causal inference and debiasing are critical tech supports for building a fair, diverse recommendation ecosystem, ensuring the long-term health of the platform.
LLMs for Recommendation
LLM Integration Trends:
- Foundation Models Revolutionizing Recommendations
- Transformer architectures brought breakthrough progress to sequence modeling, significantly enhancing behavior sequence understanding.
- Self-attention mechanisms capture connections and transitions of user interests across items.
- The pre-train and fine-tune paradigm enables recommendation models to learn from universal data before optimizing for specific scenarios.
- Core Value of LLMs in Recommendations
- Multimodal Understanding: Integrating text, images, and video to achieve rich item representations.
- Long Text Processing: Parsing item descriptions and user reviews to extract deep features.
- Cross-domain Transfer: Leveraging universal knowledge to mitigate cold start and data sparsity.
- Intent Understanding: Natural language interfaces to capture complex, vague user queries.
- Implementation Paths
- LLM as Feature Extractor: Generating high-quality item and user representations to feed into traditional models.
- Hybrid Architectures: Combining LLMs with traditional systems, where LLMs handle semantic understanding and traditional layers handle personalization.
- End-to-End LLM Recommendation: Re-defining recommendation as a generative task, directly outputting lists via LLMs.
- Retrieval-Augmented Generation (RAG): Combining RAG to boost recommendation recency and accuracy.
LLMs are reshaping the path of recommendation technologies, bringing revolutionary changes to features, architectures, and user interactions, opening up new possibilities for resolving cold-start and explainability bottlenecks.
Large-scale + Real-time + Deep Learning
Real-time Features:
- Capturing Real-time User Interests
- User interests shift rapidly; real-time features capture the latest preferences.
- Short-term interests often dictate current actions more than long-term interests.
- Real-time features significantly boost relevance and timeliness.
- Meeting Latency Demands
- News and short video feeds require extreme real-time updates.
- E-commerce promotions and live streams must adapt to real-time user actions.
- Real-time features reduce recommendation delays, boosting user experience.
- Challenges in Real-time Feature Engineering
- Real-time collection and processing of massive user behavior logs.
- Balancing calculation efficiency and storage costs.
- Maintaining consistency between real-time and offline features.
Real-time Models:
- Necessity of Real-time Updates
- Content distribution environments change quickly; models must adapt.
- Cold-start items need fast, accurate model evaluations.
- Real-time updates prevent model drift, maintaining high performance.
- Technical Challenges in Online Learning
- Balancing incremental updates with full-batch training.
- Balancing update frequencies with computing resource constraints.
- Co-optimizing real-time features and model parameters.
- Real-time Architecture Design
- Deploying stream computing frameworks (Flink, Spark Streaming).
- Storage selections (Redis, Cassandra) for real-time reads.
- Hybrid online and nearline learning architectures.
Strategies for Boosting Real-time Capacity:
- Engineering Solutions
- Deploying unified batch-stream feature architectures.
- Implementing incremental feature calculations and caching.
- Deploying lightweight online models for real-time prediction and heavy offline models for training.
- Algorithmic Optimizations
- Designing model structures suited for incremental updates.
- Deploying online learning algorithms (such as FTRL, Follow-the-Regularized-Leader).
- Injecting time decay factors to emphasize recent behaviors.
- Best Practices
- Building real-time feature monitoring dashboards.
- Establishing evaluation pipelines to track model performance decay over time.
- Implementing gray releases and quick-rollback capabilities for model updates.
Real-time computing is a core competitive edge of modern recommender systems, capturing user interest shifts to deliver highly relevant, timely content.