TailorMind: Towards Preference-Aligned Multimodal Content Generation Hengji Zhou1∗ , Ye Liu1∗ , Yufeng Liu1 , Si Wu1 , Lianghao Xia2† , Liqiang Nie2 1 South China University of Technology 2 Harbin Institute of Technology, Shenzhen [email protected], [email protected], [email protected], [email protected], [email protected], [email protected]
Abstract
Catalog-Bound Personalization
arXiv:2606.23643v1 [cs.AI] 22 Jun 2026
Personalized content systems depend on available UGC and struggle when suitable content is absent, delayed, or costly to create. Although multimodal generators can synthesize content on demand, how to translate behavioral traces into generation-ready preferences remains underexplored. We study personalized multimodal content generation: creating user-tailored multimodal content without existing item pools or waiting for matching UGC. We propose TailorMind, linking collaborative preference modeling with controllable multimodal generation. TailorMind enriches sparse user histories via hypergraph collaborative filtering and optimizes textual profiles with ranking-error feedback and textual gradient descent. Retrieval-augmented style control grounds outputs in authentic UGC patterns, while cross-modal cohesion reflection reduces semantic drift. We construct TailorBench, a benchmark from three mainstream platforms evaluated along five dimensions: coherence, novelty, aesthetic, hallucination, profiling. Experiments show that TailorMind achieves competitive or stronger coherence, improves novelty and aesthetic quality over representative generation baselines and ground-truth UGC, demonstrating advantages over retrieving available content or comparable UGC, while achieving up to 29% Recall gains in reranking. Our code is released at: https://github.com/ iLearn-Lab/TailorMind.
1
Delicate Fancy
Personalized AIGC
… Existing and Forthcoming UGC Delayed
Unremarkable
…
…
Personalized Multimodal Generation Personalization
Novelty
Figure 1: Personalized multimodal content generation.
cialized skills, tools, time, and production costs, making high-quality large-scale creation difficult for individual creators. Second, emerging topics depend on whether creators notice trends, collect materials, and publish in time, creating a latency gap between user interest and available content. User-aligned artificial intelligence generated content (AIGC) can complement UGC by synthesizing multimodal content on demand, helping platforms cover fresh topics, reduce creation costs, and seed early engagement in specific scenarios. Recent work increasingly uses large language models (LLMs) as user simulators for synthetic interactions (Chen et al., 2025) or ranking agents for preference modeling over item descriptions (Sun et al., 2023). Multimodal generative models also enhance item representations via cross-modal synthesis (Pan et al., 2025; Zhang et al., 2025a; Fu et al., 2024) or augmented views for contrastive learning (Ren et al., 2024). These studies show the value of generative models for user understanding, but rarely treat generation itself as the primary personalized output. Closer to our goal, PMG (Shen et al., 2024) extracts user preferences from interaction histories as explicit keywords and implicit soft embeddings, jointly conditioning a diffusion model for image generation. Pigeon (Xu et al., 2025) instead operates in the visual token space, filtering noisy history images via token-level masking and incorporating multimodal instructions. However, existing personalization mechanisms, whether catalog-bound or early generation-
Introduction
User-generated content (UGC) has become the dominant source of online media and remains central to content platforms (Nan et al., 2025; Safonov et al., 2025; Li et al., 2025). Yet relying solely on UGC leaves practical gaps. First, novel or technically demanding multimodal content requires spe1∗ Hengji Zhou and Ye Liu have equal contribution to this work. 2† Lianghao Xia is the corresponding author.
1
oriented, retain two inherent limitations. Personalization. Catalog-based methods (Zhang et al., 2024; Anand and Maurya, 2025) only select from available items, so outputs typically intersect with user tastes rather than being fully customized to unique preference profiles. Recent personalized generation methods condition generators on behavioral descriptions or history-derived visual preferences (Shen et al., 2024; Xu et al., 2025); yet without optimizing interpretable profiles against real interactions or exploiting collaborative signals to enrich sparse evidence, they struggle with finegrained, cross-modal, and platform-scale preferences. Novelty and immediacy. Catalog-based methods face a hard novelty ceiling, as recommendations are confined to existing items. When user preferences point to uncovered content—whether driven by long-tail interests (Tang et al., 2025; Zhang et al., 2025b) or unexplored preference combinations—systems can only fall back on the nearest catalog approximation. A temporal gap compounds this constraint: platforms must wait for creators to notice demand, produce, and publish, leaving emerging interests chronically underserved. Overcoming these limitations requires generating content on demand, beyond fixed catalogs and without waiting for future UGC supply, while grounding generation in behavior-derived preferences. This raises two challenges: deriving reliable, actionable profiles from sparse interactions and preserving them throughout multimodal generation.
content generation from user preferences. For C1, TailorMind augments sparse histories with hypergraph collaborative filtering and refines textual profiles through textual gradient descent driven by ranking errors on real interactions. For C2, retrieval-augmented style control grounds generation in user-aligned exemplars, while cross-modal cohesion monitoring corrects semantic drift between text and visual modalities. These components generate content that reflects individual preferences while preserving stylistic authenticity. Validating personalized generation requires benchmarks with rich multimodal content, real user interactions, and unified evaluation across heterogeneous outputs. We introduce TailorBench, a benchmark of real-world interactions and complete multimodal content from mainstream platforms. TailorBench assesses cross-modal coherence, novelty, aesthetic, content safety, and profiling, enabling evaluation of user alignment and generative fidelity. Our contributions are summarized as follows: • We introduce personalized multimodal generation for synthesizing tailored content beyond the personalization ceiling, novelty boundaries. • We propose TailorMind, a personalized generation framework connecting augmented collaborative preference modeling with multimodal generation through data-driven profile optimization and cross-modal alignment mechanisms. • We construct TailorBench, a personalized generation benchmark with rich multimodal content from mainstream platforms and a unified fivedimensional evaluation framework.
C1. Preference-Profile Alignment. Textual profiles should capture nuanced preferences rather than generic summaries. However, sparse and noisy histories make LLM profiling prone to superficial keyword aggregation, while collaborative signals in interaction patterns are easily missed. A reliable solution should optimize profiles against user behavior, requiring feedback and update mechanisms for discrete natural-language representations.
• Comprehensive evaluations demonstrate that TailorMind achieves superior user alignment, novelty, and multimodal generation quality.
2
Personalized Multimodal Generation
Personalized multimodal generation bridges two complementary domains: user preference modeling, which infers individual interests from behavioral evidence, and multimodal content generation, which synthesizes multimodal outputs from creative instructions and references. Task Definition. Let U = {u1 , . . . , uI } be the user set and V = {v1 , . . . , vJ } the item set. For each user u ∈ U, the observed interaction history is Du = {vi1 , . . . , vin }. We learn a behavioral preference representation:
C2. Aligning Generated Content with Profiles. Even with an accurate profile, multimodal generation can drift during concept planning, text writing, image synthesis, or video creation. Generated components may partially follow the profile while losing authentic UGC style or cross-modal consistency. Thus, profile constraints must be propagated and checked throughout generation to maintain stylistic authenticity and cross-modal coherence. To address these challenges, we present TailorMind, a framework for personalized multimodal
fR : (u, Du , V) → hu , 2
(1)
Multimodal Interaction Graph
Hypergraph Collaborative Filtering 𝑙=1
𝑣1
𝑒𝑖,𝑗
ℎ𝑧
ℎ𝑢′
𝑖𝑚𝑔 𝐟𝑣
𝐟𝑣𝑣𝑖𝑑
𝑅𝑢 , 𝐻𝑢
ℎ𝑣
𝑧𝑘 𝑓 𝐡𝑢 , 𝐡𝑣
text image video
ℎ𝑢
ℎ𝑣
(𝑛) 𝑃𝑢
Ranking Positive Training
Augmented Items
…𝑅
𝑢
Negative
(0)
𝑃𝑢
Preference 1: Interest in …
∇𝑡𝑒𝑥𝑡 𝐩𝑢 ℒ
CF 𝒟𝑢 𝐡𝑢
𝑅𝑢 , 𝐻𝑢
Title: Hupu Chronicles Image Content: … User Profile
Feedback
Multimodal Content Generation
⊕ Refined Profile
(𝑛)
𝑃𝑢
Encoder Semantics
…
𝑣𝑗 𝑢1
𝑋
𝐟𝑣𝒕𝒙𝒕
Textual
𝑌𝑇
𝑢𝑖
ℎ𝑢 ′
Iterative Profiling Optimization
Multimodal Info
𝑋
f𝑣
Tailored Note Post Video Idea
VDB
Retrieval
Cross-Modal
Cohesion
Figure 2: Overall architecture of the proposed TailorMind framework for personalized multimodal generation.
where hu encodes latent user preferences from historical interactions. Rather than serving as the final output, hu provides behavioral evidence for constructing explicit user profiles and creation instructions to synthesize content beyond V. Let T denote the space of textual creation instructions and M = {image, text, video} the modality set. The generation process is defined as: fG : (t, {m1 , m2 , . . .}) → O,
yielding denser connectivity than user-item bipartite graphs and reducing sparsity degradation. Formally, let Y ∈ {0, 1}|V|×|Z| denote the itemtag matrix, where Z is the tag vocabulary, and X ∈ {0, 1}|U |×|V| denote the user-item interaction matrix. We stack X with Y⊤ to construct the augmented adjacency matrix as: " # X ′ X = ∈ {0, 1}(|U |+|Z|)×|V| , (4) ⊤ Y
(2)
where t ∈ T is the creation instruction, {mi } are optional multimodal reference materials, and O∈ / V is the newly synthesized output.
which defines a unified hypergraph between the joint node set U ∪ Z and items in V. We initialize (0) (0) embeddings hu′ for each node u′ ∈ U ∪Z and hv for each item v ∈ V, then propagate collaborative information through the hypergraph: l l ′ l hl+1 = GNN h , {h : v ∈ N (u )}; Θ , (5) ′ v u u′
Benchmark. To support personalized generation research, we construct TailorBench from real-world interactions on Rednote, Bilibili, and Hupu: B = {U, V, D, M, E},
(3)
where N (u′ ) denotes the neighborhood of node u′ and Θl are trainable parameters at layer l. After L propagation layers, we rank unseen items by predicted affinity scores: L ŷuv = f hL (6) u , hv , CF(Du ; {hu }) = TopK {ŷuv }v∈V\Du . (7)
where E is a set of evaluation methods assessing five dimensions: Coherence (cross-modal consistency), Novelty (originality relative to existing UGC), Aesthetic (textual and visual quality), Hallucination (absence of fabricated information), and Profiling (quality of user profiles). All user data in TailorBench is anonymized. Detailed implementation is discussed in Sec. 4.1.
3
where f computes user-item preference scores. Top-k items augment Du with collaborative signals, providing broader preference context for profiling.
The Proposed TailorMind Framework
This section elaborates the technical details of the proposed TailorMind framework. The overall architecture is depicted in Figure 2. 3.1
3.2 User Profiling with Iterative Optimization 3.2.1 Multimodal Profile Initialization Existing profiling methods often process modalities independently or concatenate them directly, missing cross-modal semantics that reveal explicit and implicit interests. We instead distill multimodal content into textual descriptions for coherent profile generation. Given item v with multimodal inputs {cm v }m∈M , where M = {vid, img, txt}, we extract modality-specific textual features:
Hypergraph Collaborative Filtering
To generate personalized creation instructions t for user u, VLMs require preference-aligned evidence from behavioral history and multimodal content. Directly profiling from Du is often noisy and sparse, so TailorMind augments user evidence via hypergraph collaborative filtering. We treat content tags as hyperedges because they encode semantic item relations and induce multi-way connections,
fvm = ϕm (cm v ), 3
m ∈ M,
(8)
timodal features {fj }j∈C (n) of top-ranked candi-
where ϕm (·) denotes modality-specific feature extractors based on multimodal models that convert visual and textual content into natural language descriptions. Textual features from the augmented ′ evidence set Du = Du ∪ CF(Du ) are fed into the LLM to synthesize the initial profile: ′ m p(0) = LLM {f : v ∈ D }; p , (9) prof u v u
u,k
dates, and (iii) multimodal features {fg }g∈Zu of groundtruth items. The profile update is: text (n) p(n+1) = p(n) u u ⊕ ∇pu L(pu ),
where ⊕ denotes the LLM-based refinement operation incorporating the textual gradient, computed: (n) ∇text pu L = LLM pu , {fj }j∈C (n) , {fg }g∈Zu ; pgrad , u,k
where pprof guides the LLM to reason over multimodal textual features and distill them into a co(0) herent user preference profile pu , which supports subsequent personalized content generation. 3.2.2
(12)
(13) where pgrad prompts the LLM to improve profile performance. This approximates ∇pu Ev+ ∼Zu through LLM reasoning over ranking loss, analogous to policy-gradient estimation in reinforcement learning but performed in natural language space.
Textual-Gradient Profile Refinement
To move beyond heuristic initialization, we formulate profile refinement as a feedback-driven optimization process that minimizes ranking errors on the user’s validation ground-truth interaction Zu through textual gradient descent, thereby grounding profile updates in actual interactions.
Convergence Criterion. The optimization terminates when either (1) ranking loss reaches its optimum, i.e., the groundtruth item v + ranks first: rank(v + ; p(n+1) , Cu ) = 1, u
∀v + ∈ Zu ,
(14)
or (2) the process reaches the maximum iteration Nmax , ensuring efficiency across users. Thus, TailorMind shifts from heuristic-based profiling relying on prompt design to data-driven optimization grounded in groundtruth items, directly improving interaction ranking quality while preserving natural-language interpretability. Gradient optimization prompts are provided in Appendix A.5
Optimization Objective. We optimize the user profile pu to maximize ranking quality on interaction groundtruth Zu as follows: min Ev+ ∼Zu ℓrank (pu , v + ) , (10) pu ∈P
where ℓrank (pu , v + ) measures the rank-position loss of positive interaction item v + under profile pu , and P denotes the profile space of textual descriptions encoding multimodal user preferences.
3.3
Personalized Content Generation
With the optimized user profile pu , TailorMind employs specialized generation agents for content categories such as image-text posts and videos.
Ranking Loss Generation. At iteration n, we construct Cu by sampling one positive item v + ∼ Zu and M − 1 negative items, then rank them with (n) the current profile pu to obtain top-k items:
3.3.1 Retrieval-Augmented Style Control To control style, TailorMind uses retrievalaugmented generation grounded in actual UGC. Although pu captures preference semantics, directly translating it into generation instructions may miss stylistic nuances from user interactions. We retrieve top-k exemplars by computing cosine similarity between the profile embedding and multimodal feature representations fi of items from historical interactions Du and recommended candidates CF(Du ; {hu }), defined as :
Cu ={v + } ∪ Sample(V \ Zu , M − 1), (n) Cu,k = RankLLM p(n) (11) u , Cu , where M = 100 in our implementation and (n) Cu,k denotes the top-k ranked items. We compute (n)
ℓrank (pu , v + ) from whether v + appears in top-k and its rank position, yielding a discrete feedback signal that identifies profile aspects that capture user preferences or require adjustment.
Euk = arg top-k cos(pu , fi ),
(15)
i∈Du ∪CF(Du )
Textual Gradient Descent. Because profiles are textual, conventional numerical gradients are infeasible. Following agent-based optimization (Hu et al., 2024), we use a textual gradient: the LLM estimates an update direction from (i) the current (n) profile pu and ranking performance, (ii) mul-
where Euk denotes the retrieved exemplar set. These exemplars serve as few-shot references that help the generation agent emulate authentic patterns and stylistic nuances while remaining aligned with preferences encoded in pu . 4
Table 1: Statistics of the experimental datasets.
4.1
Datasets #Users #Items #Tags #Text #Images #Videos Rednote 12131 94997 8531 7.87 7.34 2.44 Hupu 23002 126300 278 6.76 11.33 1.83 Bilibili 13349 82704 8709 6.44
4.1.1
(16)
where Ψ(·) uses user profiles and selected styles to produce personalized concept prompts. To mitigate semantic drift and enforce profile-content alignment, we employ a cross-modal cohesion reflection mechanism: ClipScore measures consistency between generated visual content x̂ and paired text t̂, while Relevance measures whether t̂ remains consistent with the user’s interaction history D via cosine similarity in text embedding space: C(x̂, t̂) = ClipScore(x̂, t̂),
(17)
R(t̂, D) = Relevance(t̂, D),
(18) 4.1.2
These cohesion signals, together with pu and the (n) previous generation Gu , guide iterative reflective optimization to reduce semantic drift: G(n+1) = Φ pu , G(n) (19) u u , C(x̂, t̂), R(t̂, D)
Baseline Methods
TailorMind is compared with comprehensive baselines: i) multimodal AI generation methods: GPT-4o (Hurst et al., 2024), Gemini-2.5Flash (Team et al., 2023), Nanobanana pro (Google, 2025a), PMG (Shen et al., 2024), Pigeon (Xu et al., 2025), RAGAR (Ling et al., 2026), Veo3.1 (Google, 2025b), Sora2 (OpenAI, 2025), CIPHER (Gao et al., 2024), and PROSE (ArocaOuellette et al., 2025); ii) ID-based representation learning methods: LightGCN (He et al., 2020), BIGCF (Zhang et al., 2024), and IRLLRec (Wang et al., 2025a); and iii) LLM-based textual ranking approaches: MoRE (Qin et al., 2025), Re2LLM (Wang et al., 2025b), LLM4Rerank (Gao et al., 2025), and MLLM-MSR (Ye et al., 2025);
where Φ(·) denotes the VLM-based reflective revision function. By maximizing cross-modal cohesion, this mechanism constrains multimodal generation to remain aligned with user preferences.
4
Datasets and Evaluation Protocols
We construct TailorBench from three mainstream platforms: Rednote, featuring lifestyle short videos and graphic-text notes; Bilibili, an entertainment-oriented youth video platform; and Hupu, a sports community mainly comprising textual discussions. Table 1 reports detailed statistics. We use leave-two-out splitting for profiling, with the last interaction for testing, the penultimate one for validation, and the rest for optimization. We randomly sample 1,000 users from each dataset (Shen et al., 2024), and compare TailorMind’s generated content with baseline outputs and the ground-truth test item along five dimensions: Coherence, where image–text outputs and video-derived textual descriptions are respectively compared against the equal-weighted caption embedding of the user’s history and target item using CLIPScore (Hessel et al., 2021) and cosine similarity, reflecting personalization to some extent. Novelty, where VLMs score novelty and interestingness on a normalized [0, 1] scale (Vargas and Castells, 2011); Aesthetic, where VLMs assess textual and visual aesthetics within [0, 1] (Kao et al., 2017); and Hallucination, where web-search tools detect false or fabricated information, labeling hallucinated instances as 1 and otherwise 0. Profiling, where 1:99 negative sampling (Ding et al., 2019) and Recall@N/NDCG@N evaluate reranking performance of user performance modeling.
3.3.2 Cross-Modal Cohesion Reflection The generation framework aims to align produced multimodal content with pu . However, long-chain generation across heterogeneous modalities can introduce semantic drift, where generated components diverge from intended profile semantics as abstract preferences are translated into creative concepts and multimodal realizations. To bridge abstract profile representations and actionable creative directives, we design R = 10 prototypical idea templates covering creative motifs on contemporary social media platforms. The user profile pu is mapped to a style subset S ⊆ {1, . . . , R} to produce candidate creative ideas Iu : Iu = Ψ(pu , S)
Experimental Settings
Evaluation
We conduct extensive experiments to validate TailorMind’s effectiveness across six research questions: RQ1 compares TailorMind with representative baselines, RQ2 examines human preferences over generated content, RQ3 analyzes key modules, RQ4 studies hyperparameter sensitivity, RQ5 evaluates cost and efficiency, and RQ6 presents a qualitative case study of generated content.
4.1.3
Implementation Details
All methods are implemented following their original papers. For ID-based representation learning 5
Image-Text
Model Rednote Hupu Metric Coherence ↑ Novelty ↑ Aesthetic ↑ Hallucination ↓ Coherence ↑ Novelty ↑ Aesthetic ↑ Hallucination ↓ GPT-4o 0.6148 0.5436 0.7243 0.0110 0.6864 0.4723 0.6332 0.0911 Gemini-2.5-flash 0.6212 0.4706 0.4391 0.2787 0.6846 0.3826 0.2838 0.5305 Nanobanana pro 0.6110 0.5931 0.7262 0.0196 0.6996 0.5529 0.6739 0.0549 PMG 0.6155 0.5682 0.6732 0.0550 0.6200 0.5912 0.6572 0.2250 PIGEON 0.6050 0.6350 0.6637 0.0124 0.6280 0.6415 0.6706 0.0160 RAGAR 0.6255 0.5477 0.6420 0.0215 0.6177 0.5905 0.6584 0.0759 Groundtruth-IT 0.6175 0.3458 0.3707 0.0072 0.6147 0.2731 0.3409 0.0197 TailorMind-IT 0.6495 0.6643 0.8286 0.0053 0.7132 0.6648 0.7341 0.0100
Videos
Table 2: Generated Content Comparison on Rednote and Hupu Datasets, averaged over Three Runs.
Veo3.1 Sora2 CIPHER PROSE Groundtruth-V TailorMind-V
0.8345 0.8295 0.8040 0.8345 0.7842 0.8430
0.0000 0.0000 0.0000 0.0000 0.0000 0.0000 Ours
Table 3: Generated Content comparison on Bilibili. Novelty
Data Bilibili Metric Coherence ↑ Novelty↑ Aesthetic↑ Halluc. ↓ Veo3.1 0.7634 0.6710 0.8335 0.0000 0.7932 0.6550 0.8345 0.0000 Sora2 CIPHER 0.7800 0.6890 0.8250 0.0000 0.7913 0.6225 0.7925 0.0000 Prose Groundtruth 0.7181 0.6805 0.7802 0.0030 TailorMind 0.7947 0.6975 0.8485 0.0000
5 4 3 2 1 0
NanoBanana2
Image-Text Video
5 4 3 2 1 0
0.6735 0.6655 0.6915 0.6905 0.5985 0.7240
0.8450 0.8410 0.8330 0.8375 0.6768 0.8480
GPT-4o
Sora2
Image-Text Video
0.0000 0.0000 0.0500 0.0000 0.0000 0.0000 Veo3
Groundtruth
0.20 0.15 0.10 0.05 0.00
Image-Text Video
Figure 3: Human evaluation on created content.
idence. Higher Novelty and Quality. Tables 3 and 2 show that TailorMind generates more novel content than generation baselines and UGC while achieving the best aesthetic scores, demonstrating its ability to synthesize fresh and visually appealing user-aligned outputs across modalities. Better Coherence and Reliability. TailorMind achieves higher coherence than UGC and competitive or superior results against generation baselines. Since the coherence reference includes both user history and the target item, these gains indicate stronger history-aware alignment while validating crossmodal cohesion reflection. 4.3 Human Evaluation on Content (RQ2)
methods, we perform grid search to identify optimal hyperparameters. For long video understanding, we analyze the initial and final 8-second segments separately, then aggregate the information to generate a holistic content description. All LLMbased ranking methods and TailorMind’s profile refinement and cohesion reflection use 3 iterations. We employ ViT-B/32 for image encoding and textembedding-3-small for text encoding. User preferences are extracted by selecting the top-1 preference for personalized multimodal content generation. Our default models are Claude-Sonnet-4.5 for ranking and profiling, Gemini-2.5-Flash-Image for image generation, and Veo-3.1 for video generation. For automatic evaluation only, we use Gemini-2.5-pro as the VLM judge and GPT-5-All for web-based factual verification. 4.2
0.7228 0.7281 0.7391 0.7404 0.6231 0.7230
Hallucination
0.6675 0.6650 0.5940 0.6870 0.6731 0.6960
Aesthetic
0.7249 0.7437 0.7503 0.7429 0.7049 0.7271
To complement automatic VLM judging and reduce comparison bias with ground-truth UGC, we further conduct a human evaluation of TailorMind’s personalized generation quality. Specifically, we randomly selected 20 generated samples from each image-text and video category and asked 18 volunteers to evaluate anonymized and randomized outputs. Each sample was rated on a 5-point scale for Novelty and Aesthetic quality, while Hallucination was labeled as 1 if misinformation was detected and 0 otherwise; the questionnaire contained 80 items and took about 1.5 hours on average. As shown in Figure 3, TailorMind achieves stronger novelty and aesthetic quality than generation baselines and ground-truth UGC, while maintaining near-human hallucination performance, validating
Overall Performance Comparison (RQ1)
We compare TailorMind with baselines from multiple perspectives; results are summarized in Tables 4, 3, and 2. Stronger User Profiling. For the profiling dimension, generated user profiles rerank 1:99 candidate sets, and Recall/NDCG measure whether they recover held-out interactions. Table 4 shows that TailorMind consistently improves Recall across all datasets, indicating that textual-gradient profile optimization better captures fine-grained preferences from sparse behavioral ev6
Table 4: Reranking performance comparison, in terms of Recall and NDCG. Boldface marks the best reranking result. Standard deviations are computed over five runs. Bilibili Rednote Hupu R@10 N@10 R@20 N@20 R@10 N@10 R@20 N@20 R@10 N@10 R@20 N@20 0.3820 0.2554 0.4780 0.2796 0.4460 0.3605 0.4980 0.3736 0.1160 0.0532 0.2190 0.0791 0.3100 0.2100 0.4120 0.2357 0.4190 0.3575 0.4600 0.3678 0.1130 0.0557 0.2210 0.0825 0.4080 0.2376 0.5240 0.2674 0.4780 0.3446 0.5400 0.3605 0.3320 0.1780 0.4800 0.2155 0.1020 0.0490 0.2050 0.0746 0.4390 0.3172 0.5250 0.3436 0.3010 0.1243 0.4740 0.1678 0.3972 0.2633 0.5260 0.2957 0.5460 0.3979 0.6400 0.4216 0.4320 0.2100 0.6040 0.2532 0.3360 0.2265 0.4470 0.2548 0.4590 0.3415 0.5430 0.3628 0.3870 0.1586 0.5710 0.1894 0.3820 0.2333 0.5150 0.2665 0.5530 0.3812 0.6420 0.4052 0.4170 0.2085 0.5960 0.2420 0.5260 0.2643 0.6770 0.3023 0.6490 0.4140 0.6800 0.4310 0.4500 0.1919 0.6230 0.2311 TailorMind ±8.2e-4 ±1.5e-2 ±2.0e-2 ±1.3e-2 ±1.6e-2 ±1.2e-2 ±1.3e-2 ±1.1e-2 ±1.1e-2 ±5.1e-3 ±1.4e-2 ±5.4e-3
Data Metrics LightGCN BIGCF IRLLRec MoRE Re2LLM LLM4Rerank MLLM-MSR