ConceptioArchivearXiv CS
arXiv CSopen access

Learning Color Grading, No Photo Sharing: Federated Aesthetic Preference Learning for Personalized Image Enhancement

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
clouddistributedcomputingparallelcomputing
distributed computing, parallel computing, cloud

Learning Color Grading, No Photo Sharing: Federated Aesthetic Preference Learning for Personalized Image Enhancement Chuanzhi Xu1∗† , Ziyuan Tao2∗ , Jean Julien KNell1 , Yanrong Chen1 , Haolan Guo1 , Xuanhua Yin1 , Adnan Mahmood2 , Weidong Cai1 1

2 The University of Sydney Macquarie University [email protected] § Project Page

Personalized image enhancement should reflect individual aesthetic taste, yet learning such preferences commonly depends on private photos and ratings that are unsuitable for centralized collection. The task must infer preference from sparse, heterogeneous feedback and translate it into naturallooking color transformations on resource-constrained user devices. We introduce FedPAIE, a federated personalized aesthetic image enhancement framework for user-adaptive color grading without centralizing raw photos or ratings. FedPAIE trains a lightweight dual-cue aesthetic scorer, calibrates it into a personalized scorer on a small local support set, and freezes it to guide regularized adaptation of a lightweight CLUT enhancer from unpaired local photographs. Fidelity constraints and an excess-gap penalty regularize scorer-guided adaptation to limit proxy-score over-optimization while preserving content and natural appearance. Training remains lightweight throughout the pipeline: scorer learning updates at most 0.787M parameters, enhancer adaptation updates 0.265M, and inference retains only a 0.293M-parameter personalized enhancer. Experiments on MIT-Adobe FiveK and Flickr-AES demonstrate effective open-world personalization and a favorable balance between user preference and image fidelity. FedPAIE thus connects decentralized preference learning with efficient personalized image transformation without requiring paired user retouches.

1

Introduction

Personalized aesthetic image enhancement and color grading are common in digital photography, social media, and mobile content creation. Beyond exposure and contrast correction, users expect systems to adapt to preferences for color temperature, saturation, tone curves, and stylistic mood. Platforms could learn these preferences from private photos, ratings, editing histories, or text prompts, but such sensitive data are unsuitable for centralized collection. This raises a question: how can an aesthetic enhancement model learn personalized color-grading preferences while keeping user data private? Existing image enhancement methods learn efficient, image-adaptive color and tone transformations from paired data (Bychkovsky et al. 2011; Zhang et al. 2022; Kim, Lee, and Cho 2025). Meanwhile, personalized enhancement ∗ †

These authors contributed equally. Corresponding author.

Local Scorer Update (for Global Preference Modeling)

Server

User 1

Personalized Scorer Calibration (Local Ratings) Frozen-Scorer-guided CLUT Enhancer Adaptation

Global Scorer Aggregation

arXiv:2607.27659v1 [cs.CV] 30 Jul 2026

Abstract

Local Scorer Update (for Global Preference Modeling)

User 2

Personalized Scorer Calibration (Local Ratings)

Federated Aesthetic Preference Learning Global Scorer Param. Updated Scorer Param.

Generic Enhancement Prior Learning (Generic CLUT Enhancer)

Upload Private Photos

Frozen-Scorer-guided CLUT Enhancer Adaptation Ultra-lightweight Model on User-side Device

A Large Number of Users (1~k)

Local Scorer Update (for Global Preference Modeling)

User k

⋯⋯⋯

Personalized Scorer Calibration (Local Ratings) Frozen-Scorer-guided CLUT Enhancer Adaptation

Privacy Exposure and Centralized Data Risk

✕ Users have to upload private photos ! ✕ Users will be unwilling to upload photos ! (less data for model training) ✕ Form sensitive centralized database ! ✕ Unsafe personalized preference learning ! ✕ High costs on upload and storage !

Figure 1: FedPAIE learns user-specific color grading without centralizing private photos or ratings. Federated Aesthetic Preference Learning yields a global scorer, locally calibrated into a personalized scorer and frozen to guide a lightweight 0.293M-parameter CLUT enhancer on unpaired user images.

methods model user-specific retouching preferences (Kim, Koh, and Kim 2020; Bianco et al. 2020; Kosugi and Yamasaki 2024), while personalized aesthetic assessment predicts individual deviations from population-level aesthetics (Ren et al. 2017). However, many existing approaches still assume centralized access to user data, paired user retouching examples, or preference annotations that may be sensitive in real consumer applications. These assumptions are misaligned with privacy-sensitive image-editing scenarios, where users may provide only limited feedback and may refuse to share raw photos or preference prompts, as illustrated in the lower panel of Fig. 1. Federated learning offers a data-local approach to preference learning because it trains models from decentralized user data without directly uploading raw samples to a central server (McMahan et al. 2017; Kairouz et al. 2021). Recent

work has demonstrated federated visual learning in deployed object detection, general computer-vision benchmarks, personalized aesthetic assessment, and parameter-efficient video moderation (Liu et al. 2020; He et al. 2021; Xiong, Yu, and Shen 2023; Tao et al. 2025). However, applying federated learning to personalized aesthetic enhancement remains challenging. First, federated personalization requires client-side model execution, yet many aesthetic enhancement models are too heavy for user devices (Kim, Koh, and Kim 2020; Kosugi and Yamasaki 2024). User preferences are highly subjective, while an individual user typically provides explicit ratings for only a small number of images, leaving each client with limited supervision for learning personal taste. Meanwhile, rating scales and aesthetic distributions vary substantially across users, yielding non-independent-and-identically-distributed (non-IID) federated data. Standard federated aggregation can learn a shared initialization from these decentralized signals but may not fully capture every user’s aesthetic preferences. Moreover, personalized aesthetic enhancement introduces an additional challenge: the learned preference model must be used as a training signal for image transformation. Since lightweight aesthetic scorers provide only weak and imperfect supervision, directly optimizing an enhancer against their predictions may exploit imperfections in the proxy objective (Amodei et al. 2016), motivating explicit fidelity regularization. Therefore, a practical model must balance privacy preservation, user-specific adaptation, computational efficiency, and optimization stability. In this paper, we propose FedPAIE (Federated Personalized Aesthetic Image Enhancement), to our knowledge the first federated method for personalized aesthetic image enhancement and color grading. As illustrated in Fig. 1, FedPAIE separates Federated Aesthetic Preference Learning and Generic Enhancement Prior Learning from private On-Device Preference Adaptation, keeping raw photos and ratings local while each client stage trains at most 0.787M parameters and inference retains only a 0.293M-parameter enhancer. Our contributions are summarized as follows: • We introduce FedPAIE, to our knowledge the first federated method for personalized aesthetic image enhancement and color grading, enabling unpaired enhancement while keeping user data local. • We develop the Lightweight Dual-Cue Aesthetic Scorer, Federated Aesthetic Preference Learning, Generic Enhancement Prior Learning, and Personalized Scorer Calibration. We further introduce the regularized Personalized Enhancement Objective and checkpoint selection criterion for Frozen-Scorer-Guided Enhancer Adaptation. • Extensive experiments demonstrate effective open-world personalization and a favorable preference–fidelity tradeoff. Five matched ablation settings validate the roles of key objective terms and functional groups.

2

Related Work

We briefly review related work here, with complete analysis and discussion in Appendix Sec. I.

Personalized Aesthetics Assessment and Enhancement. Generic aesthetic modeling spans population-level image assessment, aesthetics-aware diffusion generation, 3D scene assessment, etc. (Ke et al. 2023; Yin et al. 2026; Xu et al. 2026), while personalized assessment methods predict user-specific judgments from attributes, few-shot adaptation, graph collaboration, transitional contrast learning, task-vector customization, or continual feedback (Yang et al. 2022b; Zhu et al. 2022; Shi et al. 2024; Yang et al. 2024; Yun and Choo 2024; Zhong et al. 2025). These personalized assessment methods output scores rather than personalized transformations. Personalized enhancement learns user-specific retouching through preference embeddings, neural-spline transforms, masked style modeling, or global-local style conditioning (Kim, Koh, and Kim 2020; Bianco et al. 2020; Kosugi and Yamasaki 2024; Kim et al. 2025). Recent systems infer photographic styles from pairwise judgments or combine VLM-driven interaction and scene-aware memory with semantic retouching (Kim, Yoo, and Kim 2026; Chang et al. 2026). FedPAIE differs in learning from federated sparse scalar ratings, keeping raw photos and ratings local, and using a calibrated scorer for unpaired on-device enhancement without user-specific retouch targets. Federated Preference Learning. Federated learning and non-IID variants keep raw samples local during training (McMahan et al. 2017; Li et al. 2020; Mohri, Sivek, and Suresh 2019). Federated recommenders learn private preferences (Ammad-ud-din et al. 2019; Chai et al. 2021; Liang, Pan, and Ming 2021; Yi et al. 2021; Liu et al. 2023), while federated personalized image-aesthetics assessment predicts user-specific scores (Xiong, Yu, and Shen 2023). To our knowledge, FedPAIE is the first federated method for personalized aesthetic image enhancement and color grading.

3.1

3

Methodology – FedPAIE

Framework Overview

FedPAIE realizes personalized color grading with an imageadaptive 3D LUT through Generic Enhancement Prior Learning and private On-Device Preference Adaptation. As shown in Fig. 2, the pipeline has three stages. Global initialization independently learns a population-level scorer through Federated Aesthetic Preference Learning and a CLUT enhancer through Generic Enhancement Prior Learning. On-device adaptation calibrates the scorer on a private rated support set, freezes it, and then adapts the enhancer on unpaired local photographs. Personalized On-Device Inference removes the scorer and retains only the lightweight personalized enhancer. Raw photographs and ratings remain local, and no paired user-specific retouches are required. Endto-end pseudocode, masked parameter updates, and the component lifecycle appear in Appendix Secs. A, B, and D. Formally, client k owns private rated samples Dk = k {(Ii , yi )}ni=1 , where Ii is the i-th image, yi ∈ [0, 1] is its normalized aesthetic rating, and nk = |Dk |. Let θg and ϕg denote the global scorer and generic enhancer parameters, yielding Sθg and Eϕg . For a new user u, FedPAIE uses a rated support set Dus and unpaired local images Uu to obtain personalized parameters θu and ϕu . The output Iˆ = Eϕu (I)

Image

Global Scorer Param. (𝜃 ) )

User 2

(Aggregation)

Global Aesthetic Scorer (𝑆" ! )

HSV+CIE lab Statistics

24-D Color Descriptor

Color Projection

MobileNet

Semantic Embedding

Semantic Projection

Regression Objective

User Rating (𝒚𝒊 )

MIT-Adobe FiveK (Paired Retouches)

Input 𝐼

Retouch 𝐼∗

CNN Backbone

1 * 𝜔' 𝑦ˆ ' − 𝑦' ( 𝑛&

User k

'

Paired Training Objective 1 ℒ%-./0= 𝐸2 𝐼 − 𝐼∗

Coefficient Head

Input

Compressed LUT Bases

⋯⋯

Weighted LUT Fusion

Residual Color-grading Transform

Trilinear LUT Application

+

+𝜆3 𝑑45657 𝐸2 𝐼 , 𝐼∗

Retouch Target

Clip [0,1]

Enhanced Image

Private Rated Support Set

Rated Image Finetune

Local Aesthetic 🔥 Scorer (from global)

After Calibration

🔥

User

Rating Target

Masked Gradient Update

Personalized Calibration Objective 9 9 9 ℒ9: = 𝜆9#$% ℒ#$% + 𝜆99;0<# ℒ;0<# + 𝜆9=0# ℒ=0#

Frozen Personalized Scorer

Personalized Scorer (𝑆"ᵤ )

Support-Dependent Scorer Mask

CNN Backbone & Coefficient Head

Input

For each new user

Retouch 𝐼∗

Predicted Aesthetic Score (𝒚ˆ 𝒊 )

Image-adaptive Coefficients

Lightweight Coefficient Predictor Input 𝐼

& ℒ#$% =

𝜎 𝜏 𝑧'

Color Extractor

Local Scorer Param. 𝜃&)*+ & Rating Count 𝑚&

Fusion MLP

User 1

🔥

Compressed LUT Bases

Personalized Enhancement Objective ℒ91 = 𝜆;#$> ℒ;#$> + 𝜆0$? ℒ0$? + 𝜆+ ℒ+ Masked +𝜆;$#@ ℒ;$#@ + 𝜆%0; ℒ%0; Enhancer Update

Enhanced Image

🔥 Trainable Frozen

Figure 2: Overview of FedPAIE. Global initialization learns the Lightweight Dual-Cue Aesthetic Scorer and CLUT enhancer through Federated Aesthetic Preference Learning and Generic Enhancement Prior Learning. On-Device Preference Adaptation performs Personalized Scorer Calibration followed by Frozen-Scorer-Guided Enhancer Adaptation. Personalized On-Device Inference retains only the lightweight personalized enhancer. should receive a higher user-specific preference score than I while preserving its content and natural appearance.

3.2

Federated Aesthetic Preference Learning

Lightweight Dual-Cue Aesthetic Scorer. The scorer combines low-level color statistics with high-level semantic context. For image Ii , the fixed color extractor Φc computes the mean, standard deviation, minimum, and maximum of each HSV and CIE Lab channel, yielding a 24-dimensional descriptor ci = Φc (Ii ). A shared, fixed ImageNet-pretrained MobileNetV3-Large (Howard et al. 2019) serves as the semantic extractor Φs and produces a 960-dimensional embedding si = Φs (Ii ). Both feature extractors are reused by all clients and remain frozen throughout global and personalized scorer training. Trainable projections Pc and Ps , parameterized by θc and θs , map the color and semantic cues to 512and 256-dimensional latent representations: hci = Pc (ci ), hsi = Ps (si ). (1) The two latent representations are concatenated and processed by a fusion MLP Gθf parameterized by θf . Together with a learnable temperature τ , the trainable scorer parameters are θ = (θc , θs , θf , τ ). The unconstrained scalar output is mapped to the interval (0, 1) as:  ŷi = Sθ (Ii ) = σ τ Gθf [hci ∥hsi ] , (2) where τ is constrained to a fixed interval and σ is the sigmoid function. Thus, ŷi ∈ (0, 1) is predicted on the same normalized scale as yi ∈ [0, 1]. The dual-cue design keeps the trainable scorer lightweight while retaining the color sensitivity needed for grading and the semantic context needed to assess whether a transformation suits the image.

Federated Global Preference Modeling. The federated stage learns a population-level scorer initialization using rating regression only. Pairwise ordering and variance preservation are introduced during user-specific calibration. For client k, the regression objective is: 1 X ωi (ŷi − yi )2 , (3) Lkreg = nk i where ωi = 1 by default. When a client’s rating distribution exceeds a prespecified imbalance threshold, ωi becomes the normalized and clipped inverse frequency of the rating bin containing yi . This optional reweighting changes individual error contributions without adding an objective. Appendix Sec. B gives the binning, activation, normalization, and clipping rules. The global client objective is therefore: LkFL = Lkreg .

(4)

Let θ denote the server-side scorer parameters at the start of communication round t, with θ0 denoting their initialization, and let Ct be the participating clients. The server sends θt to each k ∈ Ct . After minimizing Equation (4), client k returns only its updated parameters θkt+1 and the count mk of examples processed across its local steps. No raw sample or rating is transmitted. We use a square-root-reweighted variant of FedAvg (McMahan et al. 2017): √ X mk θt+1 = αk θkt+1 , αk = P √ . (5) mj j∈Ct t

k∈Ct

Here αk is client k’s normalized aggregation weight. Squareroot weighting retains a notion of client evidence while

Person Enhanc

reducing domination by users with many more ratings. FedPAIE establishes protocol-level raw-data locality: photographs and ratings remain on client devices, while federated communication contains only lightweight scorer parameters and aggregation counts. Secure aggregation and differential privacy are compatible communication-layer extensions, as detailed in Appendix Secs. D and H.

3.3

Generic Enhancement Prior Learning

We use CLUT-Net (Zhang et al. 2022) as the generic enhancer because it represents color grading as an efficient image-adaptive transform. Let ϕ denote its complete parameter set, including a lightweight coefficient predictor and compressed LUT bases. Given I, the coefficient predictor Wϕ , comprising a CNN backbone and coefficient head, produces image-adaptive coefficients w(I) = [w1 (I), . . . , wM (I)]. Weighted fusion combines M compressed LUT bases {Bq }M q=1 into an image-specific residual LUT: M X Lϕ (I) = wq (I)Bq . (6) q=1

The compressed bases are reconstructed from factorized parameters. Trilinear LUT application T evaluates the fused LUT at the input RGB values to produce a residual colorgrading transform. Appendix Sec. C gives the factor dimensions, basis reconstruction, and frozen-basis personalization rule. Adding this residual to the input gives the unclipped output:  eϕ (I) = I + T I, Lϕ (I) . E (7) During Frozen-Scorer-Guided Enhancer Adaptation and Personalized On-Device Inference, an external clipping step produces the valid image:   eϕ (I), 0, 1 . (8) Eϕ (I) = clip E The residual formulation preserves the input as a natural reference and enables full-resolution enhancement without a heavy pixel-generating decoder. FedPAIE initializes ϕg from a CLUT-Net checkpoint pretrained on paired MIT-Adobe FiveK retouching data. For a pair (I, I ∗ ), where I ∗ is an expert-retouched target, the Paired Training Objective evaluates pixel fidelity and LPIPS (Zhang eϕ (I): et al. 2018) on the unclipped output E h i ∗ e LE global = E(I,I ∗ ) ∥Eϕ (I) − I ∥1 h i (9) eϕ (I), I ∗ ) . + λp E(I,I ∗ ) dLPIPS (E Here λp ≥ 0 weights the perceptual term. This initialization supplies diverse, generally useful color transformations before any private preference signal is introduced.

3.4

On-Device Preference Adaptation

Personalized Scorer Calibration. For user u, we initialize the local scorer from θg and apply the Support-Dependent Scorer Mask MSu to select the parameter blocks adapted on the private rated support set Dus . The 10-shot regime updates

only (θf , τ ). The 100-shot regime also updates (θc , θs ), while both feature extractors remain fixed. This policy restricts capacity under sparse supervision and permits stronger feature alignment when more ratings are available. Appendix Sec. B gives the exact binary masks and their relation to the implementation cutoff. Thus, 10-shot calibration adjusts the fusion and rating scale without relearning cue projections, while 100-shot calibration can realign both projected cues. This explicitly controls adaptation capacity instead of fine-tuning the full scorer from sparse ratings. Beyond the global regression objective, local calibration models relative preferences and discourages prediction collapse. For a local mini-batch, let δ > 0 be the minimum normalized-rating separation, let Pu = {(i, j) : i < j, |yi −yj | > δ} contain the resulting unordered pairs, and let rij = sign(yi − yj ). We use the smooth pairwise objective: Lupair = −

1 |Pu |

X

log σ(rij (ŷi − ŷj )) .

(10)

(i,j)∈Pu

Let y and ŷ collect the target and predicted ratings in the same mini-batch. We also use the batchwise variance-preservation term: Luvar = [ρ Std(y) − Std(ŷ)]+ ,

[z]+ = max(z, 0), (11) where ρ ∈ (0, 1] specifies the fraction of target-score dispersion to preserve. Let Lureg denote Equation (3) evaluated on Dus . The weighted regression, pairwise, and variance terms form the personalized calibration objective LSu in Fig. 2. We collect the adaptable scorer parameters in ϑ, with feasible set Au , and hold all remaining parameters at their global values g θfix . Calibration solves:   eu Lu + λu Lu , ϑu = arg min λureg Lureg + λ pair pair var var ϑ∈Au

g θu = (θfix , ϑu ). (12)

eu , and λu weight reThe nonnegative coefficients λureg , λ var pair gression, pairwise ordering, and variance preservation. The effective pairwise coefficient incorporates support-regime selection and optional collapse protection, as specified in Appendix Sec. B. If the support set does not contain reliable ordered pairs, the pairwise term is omitted. After calibration, Sθu is frozen and used only as a training-time preference model. Frozen-Scorer-Guided Enhancer Adaptation. We initialize a local CLUT enhancer from ϕg and freeze both its compressed LUT bases and the personalized scorer. Only the CNN backbone and coefficient head of the lightweight coefficient predictor are updated. Freezing the bases preserves the transformation dictionary learned from paired retouches, while updating the predictor personalizes how those transformations are mixed rather than learning unconstrained LUTs from sparse unpaired data. For an unpaired local image I ∈ Uu , let Iˆ = Eϕ (I) and define: ˆ s+ u = Sθu (I),

s0u = Sθu (I),

0 ∆u = s+ u −su . (13)

Input

Global

Client 104

Client 108

Client 123

Client 153

Client 172

Figure 3: Five shared-input FiveK comparisons. The CLUT enhancer obtained through Generic Enhancement Prior Learning and five personalized enhancers produce distinct temperature, exposure, and contrast while preserving scene structure. Appendix Sec. G provides detail crops and full-cohort analyses. The scorer parameters remain fixed, but gradients propagate through Sθu to the enhancer. This separates a stable preference objective from the transformation being optimized and prevents joint scorer–enhancer drift. Appendix Sec. B gives the masked update and gradient derivation. All expectations below are over I ∼ Uu . The preference terms are: Laes = −EI s+ u.

(14)

ˆ I). Lperc = EI dLPIPS (I,

(15)

Lpref = −EI log σ(∆u ), The fidelity terms are: L1 = EI ∥Iˆ − I∥1 , The excess-gap penalty is:

Lgap = EI [∆u − µ]+ ,

(16)

where µ ≥ 0 is the tolerated gain. This term discourages scorer-proxy exploitation. The Personalized Enhancement Objective is: LE u = λpref Lpref + λaes Laes + λ1 L1 + λperc Lperc + λgap Lgap .

(17)

The nonnegative weights balance relative and absolute preference improvement, image fidelity, and proxy-exploitation control. The preference and aesthetic terms encourage relative and absolute improvement, the pixel and perceptual terms preserve content and natural appearance, and the excess-gap penalty limits exploitation of an imperfect scorer. Training minimizes Equation (17) over the feasible parameter set Φu , which contains only the coefficient predictor’s CNN backbone and head while keeping the LUT bases fixed. The training objective regularizes individual outputs, while model selection controls the checkpoint-level

preference–fidelity trade-off. To avoid selecting a checkpoint solely for a high proxy score, let Vu be the local validation subset and Hu the candidate checkpoint set. We score each ϕ ∈ Hu as: Qu (ϕ) = EI∈Vu [∆u ] − γ1 EI∈Vu ∥Eϕ (I) − I∥1 − γp EI∈Vu dLPIPS (Eϕ (I), I).

(18)

Here γ1 , γp ≥ 0 weight pixel and perceptual deviations, respectively. The fixed-hyperparameter (fixed-HP) configuration sets γ1 = γp = 0 and therefore selects the checkpoint with the largest validation preference gain. Shared enhancer hyperparameter optimization (HPO) uses positive fidelity penalties and applies the full regularized criterion. The final personalized parameters are: ϕu = arg max Qu (ϕ). ϕ∈Hu

(19)

Positive γ1 and γp favor user-specific improvement while rejecting checkpoints that obtain it through excessive visual deviation. The fixed-HP configuration supplies a prespecified preference-gain reference, while shared enhancer HPO exposes the attainable proxy-preference–fidelity tradeoff. Both configurations learn ϕu from ordinary local photographs without paired personalized retouching targets.

3.5

Personalized On-Device Inference

For a new photograph, deployment retains only the 0.293Mparameter enhancer Eϕu , removing the personalized scorer and its feature extractors after adaptation. In a single forward pass, the lightweight coefficient predictor estimates image-adaptive coefficients, reconstructs and combines the compressed shared LUT bases, and applies the resulting

Support Scorer initialization

SRCC↑ PLCC↑ MSE↓

10 10 10

Centralized 0.5291 Federated, fixed HP 0.5412 Federated, per-user HPO 0.5413

0.5404 0.0532 0.5471 0.0589 0.5375 0.0700

100 100 100

Centralized 0.5690 Federated, fixed HP 0.5649 Federated, per-user HPO 0.5625

0.5727 0.0552 0.5665 0.0595 0.5646 0.0609

Table 1: Test performance after Personalized Scorer Calibration on unseen users. Per-user scorer HPO fits the support set and selects on disjoint validation data. image-wide personalized color-grading transform. This path requires neither scorer evaluation nor per-image optimization, user feedback, or federated communication. Its compact single-pass design supports mobile and real-time deployment potential. Appendix Secs. E.4 and H give the resource accounting and benchmark scope.

4

Experiments and Results

4.1

Experimental Setup

4.2

Main Results

FedPAIE uses a pretrained CLUT-Net checkpoint trained on MIT-Adobe FiveK (Bychkovsky et al. 2011) as the common CLUT enhancer initialization supplied by Generic Enhancement Prior Learning. FiveK contains 5,000 images and five expert retouches. Flickr-AES (Ren et al. 2017) contains approximately 40,000 images rated by 210 users. We reserve 37 users for open-world evaluation. Filtering the remainder leaves 87 clients for 20 federated rounds and 10 fixed global-validation identities. Each user follows a fixed 70/10/10/10 training, personalization, validation, and test split. Shot counts include only support ratings, and enhancers require validation SRCC of at least 0.10. Appendix Sec. E provides the remaining protocol details. Scoring uses MSE, SRCC, and PLCC. Enhancement uses ˆ − Sθ (I) the scorer-predicted preference gain ∆u = Sθu (I) u as an optimization-aligned personalization proxy. PSNR, SSIM, and LPIPS (Zhang et al. 2018) independently measure input preservation or similarity to the non-personalized Expert C reference (the FiveK evaluation ground truth). Appendix Secs. E.5 and H give the metric definitions and interpretation scope. We first evaluate Federated Aesthetic Preference Learning and Personalized Scorer Calibration. Round 13 attains the peak validation SRCC of 0.5623 and is selected for personalization. Appendix Sec. F reports the complete global trajectory, including loss, MSE, PLCC, and prediction spread. Tab. 1 shows that the federated initialization with fixed calibration hyperparameters gives the strongest 10-shot PLCC and nearly ties per-user HPO in SRCC, suggesting that Federated Global Preference Modeling supplies a useful prior for sparse Personalized Scorer Calibration. At 100 shots, additional ratings reduce sensitivity to initialization. Fixed calibration remains competitive in scorer-only metrics. We retain per-user scorer HPO downstream because Appendix Sec. G shows higher corresponding frozen-scorer outputs

Support / Enhancer

Scorer proxy

Flickr-AES input reference FiveK Expert C reference

Score↑

PSNR↑ SSIM↑ LPIPS↓ PSNR↑ SSIM↑ LPIPS↓

∆↑

10 / Fixed HP 0.5039t 0.0247v 10 / HPO† 0.5624t 0.0826t 100 / Fixed HP 0.5245t 0.0244v 0.5864t 0.0858t 100 / HPO†

30.43 20.51 31.12 19.76

0.9627 0.7665 0.9716 0.7419

0.0162 0.1206 0.0132 0.1348

17.97 18.33 18.12 18.58

0.801 0.777 0.804 0.775

0.130 0.147 0.129 0.147

Table 2: Personalized enhancement trade-off. Input-reference metrics use Flickr-AES, while metrics relative to Expert C use FiveK. † Shared enhancer HPO selects on validation partitions from eight evaluation identities without using test images. Superscripts v/t mark validation/test proxy quantities. Proxy and fidelity axes are distinct. and input-reference PSNR across both enhancer configurations and support regimes. Appendix Sec. F provides userlevel variation and convergence analysis. We next evaluate Frozen-Scorer-Guided Enhancer Adaptation. Tab. 2 contrasts the prespecified fixed-HP open-world reference with shared enhancer HPO for within-cohort trade-off and matched ablation analyses. Both use the same per-user-HPO scorers and differ only in Frozen-Scorer-Guided Enhancer Adaptation. Further details are provided in Appendix Secs. E and F. We include SpliNet (Bianco et al. 2020), PieNet (Kim, Koh, and Kim 2020), and Masked Style Modeling (Kosugi and Yamasaki 2024) as literature context, and evaluate a controlled personalized AdaInt baseline under our shared protocol (Yang et al. 2022a). These rows provide literature context or same-protocol evidence. Appendix Sec. I distinguishes newer task settings that are not directly comparable. The fixed-HP configuration emphasizes input fidelity, while shared enhancer HPO explores stronger transformations under the regularized selection criterion. Against Expert C, shared enhancer HPO yields higher PSNR, whereas fixed HP retains higher SSIM and lower LPIPS, revealing complementary proxy–fidelity operating points. Fig. 3 shows cooler, neutral, and darker client-dependent transforms without spatial modification, as expected from user-specific color grading. Fig. 4 adds four 5/5 examples as user 204’s preference context alongside 100-shot personalized outputs. Appendix Sec. G gives extended analyses. Under the shared Expert C protocol in Tab. 3, FedPAIE raises SSIM from 0.756 to 0.801/0.804 and reduces LPIPS from 0.231 to 0.130/0.129 for 10/100 shots. AdaInt retains 1.48/1.33 dB higher PSNR, Method

Train./src. PSNR↑ SSIM↑ LPIPS↓ Params.

SpliNet PieNet Masked Style Model.

C/R C/R C/R

18.74 20.52 22.98

0.819 0.850 0.897

Original input AdaInt + pers. scorer FedPAIE, 10-shot FedPAIE, 100-shot

–/S C+L/S F+L/S F+L/S

17.84 19.45 17.97 18.12

0.791 0.756 0.801 0.804

– – –

0.03M 28M 90M

0.138 – 0.231 ∼0.6M 0.130 0.293M 0.129 0.293M

Table 3: Personalized enhancement comparison. C, F, and L denote centralized, federated, and local training, while R and S denote literature-reported and study-evaluated results. Bold indicates the best result among S rows.

(b)

(c)

(d)

Score 0.352

Reference

Score 0.571 PSNR 21.2 | SSIM 0.84

Score 0.643 PSNR 24.0 | SSIM 0.89

(e)

(f)

(g)

(h)

Score 0.633 PSNR 23.0 | SSIM 0.86

Score 0.504 PSNR 20.1 | SSIM 0.85

Score 0.632 PSNR 5.1 | SSIM 0.34

Score 0.355 PSNR 20.7 | SSIM 0.79

High Rate Image

(a)

Score 5

Score 5

Score 5

Enhanced Image

Original Image

Score 5

Figure 4: Flickr-AES user 204 (100-shot). Rows show Flickr examples rated 5 out of 5, FiveK inputs, and Full-objective outputs from Frozen-Scorer-Guided Enhancer Adaptation. The Flickr examples are unpaired preference context. 10-shot Variant

Score↑

∆↑

PC ↑ SC ↑ LC ↓ Score↑

100-shot ∆↑

PC ↑ SC ↑ LC ↓

Original 0.4324 0.0000 – – – 0.4534 0.0000 – – – Generic prior 0.4939 +0.0615 22.60 0.904 0.087 0.5153 +0.0619 22.60 0.904 0.087 Full objective 0.5213 +0.0890 18.33 0.777 0.147 0.5438 +0.0904 18.58 0.775 0.147 −Lpref 0.5160 +0.0836 18.46 0.789 0.139 0.5436 +0.0903 18.45 0.775 0.145 0.5295 +0.0971 13.43 0.495 0.341 0.5513 +0.0979 14.23 0.534 0.310 −Lgap − reg. group 0.5173 +0.0850 7.30 0.235 0.548 0.5269 +0.0736 7.73 0.289 0.529 − scorer guid. 0.4326 +0.0002 18.05 0.814 0.122 0.4535 +0.0001 18.04 0.814 0.122

Table 4: Objective ablation for Frozen-Scorer-Guided Enhancer Adaptation. Score is the mean frozen personalized scorer output, and ∆ is its change from Original. PC , SC , and LC denote PSNR, SSIM, and LPIPS relative to Expert C. “− reg. group” removes L1 , Lperc , and Lgap . Generic prior denotes the CLUT enhancer obtained through Generic Enhancement Prior Learning. Proxy and fidelity axes jointly characterize the preference–fidelity trade-off. showing complementary pixel- and perceptual-fidelity operating points. FedPAIE also improves all three metrics over the input and uniquely combines Federated Aesthetic Preference Learning with Frozen-Scorer-Guided Enhancer Adaptation. Its 0.293M-parameter enhancer uses roughly half as many parameters as the study-evaluated personalized AdaInt baseline and about 1/96 and 1/307 as many as PieNet and Masked Style Modeling, respectively. Literaturesourced rows broaden the quality–model-scale context, while the study-evaluated rows provide the same-protocol evidence detailed in Appendix Sec. F. Training and Personalized On-Device Inference remain lightweight. Federated Aesthetic Preference Learning, 10shot Personalized Scorer Calibration, and Frozen-ScorerGuided Enhancer Adaptation update at most 0.787M, 0.527M, and 0.265M parameters, respectively. Static costs are 4.70 MFLOPs per scorer sample and 5.16 GFLOPs per 224 × 224 enhancer image. Personalized On-Device Inference retains only the 0.293M-parameter enhancer. Appendix

Figure 5: Qualitative objective ablation for user 204 (100-shot): (a) Original, (b) Expert C reference, (c) Generic Enhancement Prior, (d) Full objective, (e) w/o the preference-ranking loss Lpref , (f) w/o the excess-gap penalty Lgap , (g) w/o the fidelity-plus-gap regularization group (L1 , Lperc , Lgap ), and (h) w/o frozen-scorer guidance (Lpref , Laes , Lgap ). With matched settings, Full improves both the frozen-scorer output and Expert C fidelity over the Generic Enhancement Prior. Sec. E.4 gives complete resource accounting, while Appendix Sec. G provides the image-conditioned variation and image-suitability analyses.

4.3

Ablation Study

All Frozen-Scorer-Guided Enhancer Adaptation variants share the data, frozen personalized scorers, eligibility rule, optimization protocol, and the shared enhancer HPO configuration selected for the Full objective. Holding them fixed isolates the active terms in Equation (17) without variantspecific search. These matched runs test objective-term contributions at a common within-cohort operating point, while fixed HP supplies the open-world reference. Appendix Sec. G gives the full protocol. Tab. 4 and Fig. 5 compare the Full objective with matched removal variants. Scorer guidance drives the scorer-predicted preference gain, while Lgap and the fidelity-plus-gap regularization group protect fidelity. Complete results appear in Appendix Sec. G. Under the frozen personalized-scorer proxy, Full exceeds the Generic Enhancement Prior for every evaluated user in both support regimes, with paired tests yielding p ≤ 3.6 × 10−6 .

5

Conclusion

We presented FedPAIE, a federated framework for personalized aesthetic image enhancement that keeps raw photos and ratings local. Federated preference learning initializes a scorer, which is calibrated and frozen to guide unpaired local CLUT adaptation. The pipeline updates at most 0.787M scorer and 0.265M enhancer parameters and retains only the 0.293M enhancer at inference. Experiments demonstrate effective 10- and 100-shot personalization. Controlled ablations identify scorer guidance as the preference signal and fidelity-plus-gap regularization as the mechanism balancing preference gain with image fidelity. FedPAIE thus provides a lightweight path from decentralized aesthetic feedback to personalized image transformation.

References

Ammad-ud-din, M.; Ivannikova, E.; Khan, S. A.; Oyomno, W.; Fu, Q.; Tan, K. E.; and Flanagan, A. 2019. Federated Collaborative Filtering for Privacy-Preserving Personalized Recommendation System. arXiv:1901.09888. Amodei, D.; Olah, C.; Steinhardt, J.; Christiano, P.; Schulman, J.; and Mané, D. 2016. Concrete Problems in AI Safety. arXiv:1606.06565. Bianco, S.; Cusano, C.; Piccoli, F.; and Schettini, R. 2020. Personalized Image Enhancement Using Neural Spline Color Transforms. IEEE Transactions on Image Processing, 29: 6223–6236. Bychkovsky, V.; Paris, S.; Chan, E.; and Durand, F. 2011. Learning Photographic Global Tonal Adjustment with a Database of Input/Output Image Pairs. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 97–104. Chai, D.; Wang, L.; Chen, K.; and Yang, Q. 2021. Secure Federated Matrix Factorization. IEEE Intelligent Systems, 36(5): 11–20. Chang, Z.; Duan, Z.-P.; Zhang, J.; Guo, C.-L.; Liu, S.; Chun, H.; Park, H.; Liu, Z.; and Li, C. 2026. PerTouch: VLMDriven Agent for Personalized and Semantic Image Retouching. Proceedings of the AAAI Conference on Artificial Intelligence, 40(4): 2752–2759. Chen, Y.-S.; Wang, Y.-C.; Kao, M.-H.; and Chuang, Y.-Y. 2018. Deep Photo Enhancer: Unpaired Learning for Image Enhancement from Photographs with GANs. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 6306–6314. Christiano, P. F.; Leike, J.; Brown, T. B.; Martic, M.; Legg, S.; and Amodei, D. 2017. Deep Reinforcement Learning from Human Preferences. In Advances in Neural Information Processing Systems, volume 30, 4299–4307. Deng, Y.; Loy, C. C.; and Tang, X. 2018. Aesthetic-Driven Image Enhancement by Adversarial Learning. In Proceedings of the 26th ACM International Conference on Multimedia, 870–878. Gharbi, M.; Chen, J.; Barron, J. T.; Hasinoff, S. W.; and Durand, F. 2017. Deep Bilateral Learning for Real-Time Image Enhancement. ACM Transactions on Graphics, 36(4): 118:1–118:12. He, C.; Shah, A. D.; Tang, Z.; Fan, D.; Sivashunmugam, A. N.; Bhogaraju, K.; Shimpi, M.; Shen, L.; Chu, X.; Soltanolkotabi, M.; and Avestimehr, S. 2021. FedCV: A Federated Learning Framework for Diverse Computer Vision Tasks. arXiv:2111.11066. Howard, A.; Sandler, M.; Chu, G.; Chen, L.-C.; Chen, B.; Tan, M.; Wang, W.; Zhu, Y.; Pang, R.; Vasudevan, V.; Le, Q. V.; and Adam, H. 2019. Searching for MobileNetV3. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 1314–1324. Kairouz, P.; McMahan, H. B.; Avent, B.; Bellet, A.; Bennis, M.; Bhagoji, A. N.; Bonawitz, K. A.; Charles, Z.; Cormode, G.; Cummings, R.; et al. 2021. Advances and Open Problems in Federated Learning. Foundations and Trends in Machine Learning, 14(1–2): 1–210.

Kang, S. B.; Kapoor, A.; and Lischinski, D. 2010. Personalization of Image Enhancement. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 1799–1806. Ke, J.; Ye, K.; Yu, J.; Wu, Y.; Milanfar, P.; and Yang, F. 2023. VILA: Learning Image Aesthetics from User Comments with Vision-Language Pretraining. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 10041–10051. Kim, H.-U.; Koh, Y. J.; and Kim, C.-S. 2020. PieNet: Personalized Image Enhancement Network. In Computer Vision – ECCV 2020, volume 12375 of Lecture Notes in Computer Science, 374–390. Springer. Kim, J.; Yoo, J.; and Kim, S. J. 2026. Learning Personalized Photographic Style from Pairwise User Preferences. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 1134–1144. Kim, J.-S.; Woo, S.; Kim, H.; and Kim, C.-S. 2025. Personalized Image Enhancement Using Global and Local Style Information. In 2025 International Technical Conference on Circuits/Systems, Computers, and Communications (ITCCSCC), 1–6. IEEE. Kim, W.; Lee, K.; and Cho, N. I. 2025. Lightweight and Fast Real-time Image Enhancement via Decomposition of the Spatial-aware Lookup Tables. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 11895–11905. Kosugi, S.; and Yamasaki, T. 2024. Personalized Image Enhancement Featuring Masked Style Modeling. IEEE Transactions on Circuits and Systems for Video Technology, 34(1): 140–152. Li, T.; Sahu, A. K.; Zaheer, M.; Sanjabi, M.; Talwalkar, A.; and Smith, V. 2020. Federated Optimization in Heterogeneous Networks. In Proceedings of Machine Learning and Systems, volume 2, 429–450. Liang, F.; Pan, W.; and Ming, Z. 2021. FedRec++: Lossless Federated Recommendation with Explicit Feedback. Proceedings of the AAAI Conference on Artificial Intelligence, 35(5): 4224–4231. Liu, W.; Chen, C.; Liao, X.; Hu, M.; Yin, J.; Tan, Y.; and Zheng, L. 2023. Federated Probabilistic Preference Distribution Modelling with Compactness Co-Clustering for PrivacyPreserving Multi-Domain Recommendation. In Proceedings of the Thirty-Second International Joint Conference on Artificial Intelligence, 2206–2214. Liu, Y.; Huang, A.; Luo, Y.; Huang, H.; Liu, Y.; Chen, Y.; Feng, L.; Chen, T.; Yu, H.; and Yang, Q. 2020. FedVision: An Online Visual Object Detection Platform Powered by Federated Learning. Proceedings of the AAAI Conference on Artificial Intelligence, 34(8): 13172–13179. Lv, P.; Fan, J.; Nie, X.; Dong, W.; Jiang, X.; Zhou, B.; Xu, M.; and Xu, C. 2023. User-Guided Personalized Image Aesthetic Assessment Based on Deep Reinforcement Learning. IEEE Transactions on Multimedia, 25: 736–749. Maerten, A.-S.; Chen, L.-W.; De Winter, S.; Bossens, C.; and Wagemans, J. 2025. LAPIS: A Novel Dataset for Personalized Image Aesthetic Assessment. In Proceedings of

the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, 6292–6301. McMahan, H. B.; Moore, E.; Ramage, D.; Hampson, S.; and Agüera y Arcas, B. 2017. Communication-Efficient Learning of Deep Networks from Decentralized Data. In Proceedings of the 20th International Conference on Artificial Intelligence and Statistics, volume 54 of Proceedings of Machine Learning Research, 1273–1282. PMLR. Mohri, M.; Sivek, G.; and Suresh, A. T. 2019. Agnostic Federated Learning. In Proceedings of the 36th International Conference on Machine Learning, volume 97 of Proceedings of Machine Learning Research, 4615–4625. PMLR. Ren, J.; Shen, X.; Lin, Z.; Mech, R.; and Foran, D. J. 2017. Personalized Image Aesthetics. In Proceedings of the IEEE International Conference on Computer Vision, 638–647. Shi, H.; Guo, J.; Ke, Y.; Wang, K.; Yang, S.; Qin, F.; and Chen, L. 2024. Personalized Image Aesthetics Assessment based on Graph Neural Network and Collaborative Filtering. Knowledge-Based Systems, 294: 111749. Talebi, H.; and Milanfar, P. 2018. NIMA: Neural Image Assessment. IEEE Transactions on Image Processing, 27(8): 3998–4011. Tao, Z.; Xu, C.; Jayawardana, S.; Mahmood, A.; Bao, W.; Thilakarathna, K.; and Lim, T. J. 2025. FedVideoMAE: Efficient Privacy-Preserving Federated Video Moderation. arXiv:2512.18809. Wang, G.; Yan, J.; and Qin, Z. 2018. Collaborative and Attentive Learning for Personalized Image Aesthetic Assessment. In Proceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence, 957–963. Xiong, Z.; Yu, H.; and Shen, Z. 2023. Federated Learning for Personalized Image Aesthetics Assessment. In Proceedings of the IEEE International Conference on Multimedia and Expo, 336–341. Xu, C.; Wei, B.; Zhou, H.; Yin, X.; Deng, Z.; Chen, H.; Qu, Q.; and Cai, W. 2026. Aes3D: Aesthetic Assessment in 3D Gaussian Splatting. arXiv:2605.05155. Yang, C.; Jin, M.; Jia, X.; Xu, Y.; and Chen, Y. 2022a. AdaInt: Learning Adaptive Intervals for 3D Lookup Tables on RealTime Image Enhancement. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 17522–17531. Yang, Y.; Xu, L.; Li, L.; Qie, N.; Li, Y.; Zhang, P.; and Guo, Y. 2022b. Personalized Image Aesthetics Assessment With Rich Attributes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 19861–19869. Yang, Z.; Li, L.; Yang, Y.; Li, Y.; and Lin, W. 2024. MultiLevel Transitional Contrast Learning for Personalized Image Aesthetics Assessment. IEEE Transactions on Multimedia, 26: 1944–1956. Yi, J.; Wu, F.; Wu, C.; Liu, R.; Sun, G.; and Xie, X. 2021. Efficient-FedRec: Efficient Federated Learning Framework for Privacy-Preserving News Recommendation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 2814–2824.

Yin, X.; Xu, C.; Zhou, H.; Wei, B.; and Cai, W. 2026. AccelAes: Accelerating Diffusion Transformers for Training-Free Aesthetic-Enhanced Image Generation. arXiv:2603.12575. Yun, J.; and Choo, J. 2024. Scaling Up Personalized Image Aesthetic Assessment via Task Vector Customization. In Computer Vision – ECCV 2024: 18th European Conference, Milan, Italy, September 29–October 4, 2024, Proceedings, Part XL, volume 15098 of Lecture Notes in Computer Science, 323–339. Springer. Zeng, H.; Cai, J.; Li, L.; Cao, Z.; and Zhang, L. 2022. Learning Image-Adaptive 3D Lookup Tables for High Performance Photo Enhancement in Real-Time. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(4): 2058–2073. Zhang, F.; Zeng, H.; Zhang, T.; and Zhang, L. 2022. CLUTNet: Learning Adaptively Compressed Representations of 3DLUTs for Lightweight Image Enhancement. In Proceedings of the 30th ACM International Conference on Multimedia, 6493–6501. ACM. Zhang, R.; Isola, P.; Efros, A. A.; Shechtman, E.; and Wang, O. 2018. The Unreasonable Effectiveness of Deep Features as a Perceptual Metric. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 586–595. Zhong, H.; He, S.; Ming, A.; and Ma, H. 2025. Rethinking Personalized Aesthetics Assessment: Employing Physique Aesthetics Assessment as An Exemplification. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2935–2944. Zhu, H.; Li, L.; Wu, J.; Zhao, S.; Ding, G.; and Shi, G. 2022. Personalized Image Aesthetics Assessment via MetaLearning With Bilevel Gradient Optimization. IEEE Transactions on Cybernetics, 52(3): 1798–1811.

Appendix A

End-to-End Optimization Procedure

The main paper defines the learning objectives of FedPAIE. This appendix complements those definitions with the execution order, parameter-update rules, and compressedLUT construction needed to reproduce the method. We retain the separation between two shared initializations and private adaptation. Federated Aesthetic Preference Learning and Generic Enhancement Prior Learning proceed independently, while On-Device Preference Adaptation first performs Personalized Scorer Calibration and then uses the frozen personalized scorer in Frozen-Scorer-Guided Enhancer Adaptation.

A.1

Global Initialization

Algorithm 1 summarizes global initialization. At each communication round, the server transmits only the current global scorer parameters to the selected clients. A client updates the scorer using only rating regression on its private image– rating samples and returns the updated scorer parameters. Pairwise ordering and variance preservation are not part of this federated objective. They are introduced only during Personalized Scorer Calibration. The generic enhancer is trained on a separate paired retouching corpus and therefore does not participate in the federated exchange. Client k owns k Dk = {(Ii , yi )}ni=1 , where Ii is an image, yi ∈ [0, 1] is its normalized rating, and nk = |Dk |. θk denotes the client’s current scorer parameters. At round t, θt denotes the server scorer parameters before local training, Ct the participating client set, and θkt+1 the parameters returned by client k after local training. For a local mini-batch B ⊂ Dk , the complete global meansquared-error (MSE) scorer objective is: LkFL (B) = Lkreg (B) X 1 = |B|

2 ωi Sθk (Ii ) − yi .

(20)

(Ii ,yi )∈B

The default is ωi = 1. Optional inverse-frequency weights only rebalance the regression errors. Equation (20) remains an MSE-only objective and contains no ranking or variance term.

A.2

On-Device Preference Adaptation

For a new user, Personalized Scorer Calibration must precede Frozen-Scorer-Guided Enhancer Adaptation. This ordering prevents the preference target from drifting while the enhancer is being optimized. Algorithm 2 makes the separation explicit. The Support-Dependent Scorer Mask is defined in Equation (27). The enhancer mask ME always freezes the compressed LUT bases. Checkpoints are evaluated on a local validation subset using Qu . In the fixed-hyperparameter (fixed-HP) configuration, γ1 = γp = 0, so Qu reduces to validation preference gain. The shared enhancer hyperparameter optimization (HPO) configuration uses positive fidelity penalties and therefore applies the full regularized criterion. No server interaction is required after the two shared initializations have been downloaded. The symbol ⊙ denotes

Algorithm 1: Global Initialization of FedPAIE Require: Rated sets for K clients {Dk }K k=1 , pretrained generic enhancer parameters ϕg , rounds T , local epochs Es , and learning rate ηs Ensure: Global scorer Sθg and generic enhancer Eϕg 1: Initialize scorer parameters θ0 2: for t = 0, . . . , T − 1 do 3: Sample participating clients Ct 4: for each k ∈ Ct in parallel do 5: Set θk ← θ t 6: Set processed-example count mk ← 0 7: for e = 1, . . . , Es do 8: for each local mini-batch B ⊂ Dk do 9: Compute optional regression weights {ωi }i∈B using Equation (22) 10: Evaluate Lk FL (B) using Equation (20) 11: θk ← θk − ηs ∇θk Lk FL (B) 12: mk ← mk + |B| 13: end for 14: end for 15: Return θk and the processed-example count mk 16: end for √ X mk 17: θt+1 ← θk P √ mj j∈C k∈C

18: end for 19: θg ← θT 20: return θg , ϕg

t

t

elementwise multiplication in the masked updates below. For compactness, the calibration objective in the algorithm is denoted by: eu Lu + λu Lu . LSu = λureg Lureg + λ pair pair var var

(21)

eu , and λu weight reThe nonnegative coefficients λureg , λ var pair gression, pairwise ordering, and variance preservation, respectively. The effective pairwise coefficient can be zero when the support set is too small to provide reliable ordering supervision. Its support-regime selection and optional collapse protection are specified below.

B B.1

Operational Objectives and Gradient Routing

Federated Global Preference Modeling: Regression and Rating Rebalancing

Global federated training uses Equation (20) only. To make its optional inverse-frequency weighting unambiguous, let r {Rb }B b=1 be a fixed partition of the normalized rating range, let nk,b be the number of client-k samples in bin b, and let Bk+ = {b : nk,b > 0}. Here Br is the number of bins, and b(i) denotes the bin containing rating yi . We first normalize inverse-frequency class weights by their mean over nonempty bins and then clip them for stability: 1 X nk rk,b = , r̄k = + rk,b , nk,b |Bk | + b∈Bk (22)   rk,b(i) ωi = clip , ωmin , ωmax . r̄k

Algorithm 2: On-Device Preference Adaptation for User u Require: Shared initializations (θ , ϕ paired images Uu , validation set Vu , and scorer/enhancer learning rates ηuS , ηuE Ensure: Personalized enhancer Eϕu for Personalized On-Device Inference 1: Select scorer mask MS u from the prespecified support regime 2: Initialize θ ← θg 3: for each scorer-calibration step do 4: Sample Bs ⊂ Dus and construct P(Bs ) 5: Compute Lureg , Lupair , and Luvar S 6: gθ ← MS u ⊙ ∇θ Lu (Bs ) 7: θ ← θ − ηuS gθ 8: end for 9: Set θu ← θ and freeze Sθu 10: Initialize ϕ ← ϕg and freeze the LUT factors β 11: Set best validation score Qbest ← −∞ 12: for each enhancer-adaptation epoch do 13: for each mini-batch Be ⊂ Uu do ˆ − Sθu (I) 14: Compute Iˆ = Eϕ (I) and ∆u = Sθu (I) E E 15: gϕ ← M ⊙ ∇ϕ Lu (Be ) 16: ϕ ← ϕ − ηuE gϕ 17: end for 18: Evaluate Qu (ϕ) on Vu 19: if Qu (ϕ) > Qbest then 20: Save ϕu ← ϕ and Qbest ← Qu (ϕ) 21: end if 22: end for 23: Discard Sθu 24: return Eϕu g

g

Weighting is activated only when the largest rating-bin proportion exceeds a prespecified imbalance threshold ξ. Here ξ is the activation threshold and ωmin , ωmax are the lower and upper clipping bounds. Otherwise, all ωi are set to one. Empty bins receive no samples and therefore do not contribute to the mini-batch loss. The bin boundaries, ξ, ωmin , and ωmax are implementation hyperparameters and are reported with the experimental settings. Crucially, activating these weights changes only how the squared errors are averaged. It does not add a second training signal.

B.2

ratings in Bs . The variance term is:

), rated support set Dus , un-

Personalized Pairwise and Variance Objectives

Pair construction is used only after global training, when the scorer is calibrated to a new user. For a local support s mini-batch Bs = {(Ii , yi )}B i=1 of size Bs , let δ > 0 be the minimum normalized-rating separation. We use each unordered pair once and exclude pairs whose ratings are too close to provide a reliable direction: P(Bs ) = {(i, j) : 1 ≤ i < j ≤ Bs , |yi − yj | > δ}. (23) With rij = sign(yi − yj ), the pairwise term is: X 1 log σ(rij (ŷi − ŷj )) . (24) Lupair = − |P(Bs )| (i,j)∈P(Bs )

Here σ is the sigmoid function. We set Lupair = 0 when P(Bs ) = ∅. Let y and ŷ collect the target and predicted

Luvar = [ρ Std(y) − Std(ŷ)]+ ,

(25)

where [z]+ = max(z, 0). It penalizes predictions whose dispersion falls below a fraction ρ ∈ (0, 1] of the observed rating dispersion. The operational implementation uses δ = 0.1 and ρ = 0.7 for ratings normalized to [0, 1]. To reduce the influence of unreliable ordering gradients when the scorer is close to a constant predictor, the effective pairwise coefficient can be attenuated according to:  u κλpair , Std(ŷ) < ϵc , u e λpair = (26) λupair , otherwise, where κ is the attenuation factor and ϵc the collapse threshold. The implementation uses κ = 0.5 and ϵc = 0.01. If the support-regime coefficient λupair is zero, the pairwise term remains disabled regardless of this rule. Thus, an empty-pair mini-batch still contributes through regression and variance preservation.

B.3

Masked Parameter Updates

Write the scorer parameters as θ = (θc , θs , θf , τ ), corresponding to the color projection, semantic projection, fusion head, and temperature. The pretrained MobileNetV3 semantic extractor and the deterministic color-statistics extractor are fixed and are not included in this trainable tuple. Let Nu = |Dus |. The implementation uses N0 = 20 to separate the small- and larger-support regimes. The personalization mask is:  1θf + 1τ , Nu ≤ N0 , MSu = (27) 1θc + 1θs + 1θf + 1τ , Nu > N0 , where 1a selects the coordinates of parameter block a and 0a is zero on that block. All unselected coordinates are zero. Consequently, the 10-shot configuration updates only the fusion MLP and temperature, whereas the 100-shot configuration also updates both scorer projections. The scorer update at calibration step ℓ is therefore: θℓ+1 = θℓ − ηuS MSu ⊙ ∇θ LSu ,

(28)

which guarantees that the feature extractors and every masked scorer block retain their global values exactly. Similarly, decompose the enhancer as ϕ = (ψ, β), where ψ contains the CNN backbone and coefficient head of the lightweight coefficient predictor and β contains all factorized LUT parameters. Enhancer adaptation uses: ME = (1ψ , 0β ),

β ℓ+1 = β g ,

ψ ℓ+1 = ψ ℓ − ηuE ∇ψ LE u.

(29)

Although Sθu is frozen, it remains in the differentiable path. For the preference term ℓpref = − log σ(∆u ), its gradient with respect to the trainable enhancer block is: ∇ψ ℓpref = (σ(∆u ) − 1)JE,ψ (I)⊤ ˆ · ∇ ˆSθ (I), I

(30)

u

where JE,ψ is the enhancer Jacobian. Equation (30) clarifies that the scorer supplies image-space gradients without receiving a parameter update.

C

Compressed LUT Parameterization

Let each residual LUT basis contain three color channels on a grid of resolution d. In the fully factorized form, the channel-c tensor of basis q is reconstructed from two shared factors and a basis-specific core: Bq,c (β) = reshaped×d×d (ACq,c D) ,

(31)

where rs and rw are factorization ranks, A ∈ Rd×rs , 2 Cq,c ∈ Rrs ×rw , and D ∈ Rrw ×d . The collection β = {A, D, Cq,c }q,c parameterizes all bases. Here M is the number of LUT bases. For an image I, the lightweight coefficient predictor, comprising a CNN backbone and coefficient head, produces wψ (I) ∈ RM , giving: Lψ,β (I) =

M X

Global Local S/E Deploy

Color statistics Φc Semantic extractor Φs Color projection Pc Semantic projection Ps Fusion head Gθf

F F U-FL U-FL U-FL

F/UF F/UF U∗ /UF U∗ /UF U/UF

D D D D D

U-FL U-P U-P

U/UF –/U –/F

D R R

Temperature τ Enhancer predictor ψ Compressed LUT factors β

Table 5: Lifecycle of FedPAIE components. The local column reports Personalized Scorer Calibration/Frozen-ScorerGuided Enhancer Adaptation. U, F, UF, D, and R denote updated, frozen, used but frozen, discarded, and retained. FL and P denote federated and paired global learning. * Updated only in the larger-support regime.

wψ,q (I)Bq (β),

q=1

Ie = I + T (I, Lψ,β (I)),   e 0, 1 . Iˆ = clip I,

E

(32)

T denotes trilinear interpolation of the fused LUT at the e input RGB values. The CLUT-Net forward path returns I. Its Paired Training Objective is evaluated on this unclipped output. The clipping step is applied externally during FrozenScorer-Guided Enhancer Adaptation and Personalized OnDevice Inference. The pretrained checkpoint supplies both ψ g and β g . During personalization, fixing β g preserves the learned space of plausible color transforms, while updating ψ changes how an image is mapped to a mixture of those transforms. This restriction is the architectural counterpart to the fidelity and excess-gap penalties in the Personalized Enhancement Objective.

D

Component

Component Lifecycle and Privacy Boundary

Tab. 5 collects the trainable/frozen status that is spread across the main method. Its local column reports Personalized Scorer Calibration and Frozen-Scorer-Guided Enhancer Adaptation in that order. The conditional update follows the support-size rule in Equation (27). For clarity, the server-visible state at communication round t is limited to:  t Vserver = θt , {(θkt+1 , mk ) : k ∈ Ct } . (33) The local images, ratings, constructed preference pairs, and personalized models are not uploaded. The generic enhancer is initialized from a separate generic paired corpus. After downloading (θg , ϕg ), a new user’s support, adaptation, validation, and inference stages require no further communication. This protocol enforces raw-data locality by communicating only scorer parameters and aggregation counts while keeping every user-specific asset and adaptation stage on device. The same communication boundary is compatible with secure aggregation and differential privacy when additional deployment protections are required.

E.1

Experimental Protocol and Reproducibility

Datasets and Preprocessing

MIT-Adobe FiveK (Bychkovsky et al. 2011) contains 5,000 original photographs and a retouched version from each of five experts. The preprocessing pipeline converts inputs to RGB, transforms the supplied ProPhoto RGB images to sRGB, rejects corrupted files, and resizes images while preserving orientation to either 720 × 480 or 480 × 720. For Generic Enhancement Prior Learning, FedPAIE loads a pretrained CLUT-Net checkpoint trained on paired FiveK retouching data. All personalized enhancer conditions start from this same checkpoint, so the source initialization is controlled across comparisons. Flickr-AES (Ren et al. 2017) contains approximately 40,000 images rated by 210 users on a 1–5 scale. Ratings are normalized to [0, 1]. We remove records with missing or invalid fields, duplicate user–image records, unreadable or corrupted images, and samples that cannot be converted consistently to RGB. The remaining data are organized by numeric client identifier. This dataset is unpaired: ratings support preference modeling, but there is no user-specific retouched target for enhancer personalization.

E.2

Open-World Split and Cohort Accounting

Users, rather than images alone, are separated for the global open-world protocol. Of the 210 users, 173 are candidate training users and 37 are held out from federated optimization. Each user’s records are further divided into 70% training, 10% personalization, 10% validation, and 10% test subsets with a fixed seed. The training side is filtered for at least 100 training samples and usable rating diversity, leaving 87 clients. All eligible training clients join each of the 20 communication rounds. Global validation is monitored on the held-out validation partitions of 10 fixed eligible training identities. The experiment does not subsample a new client subset from round to round. Tab. 6 separates the cohorts used by different analyses. Here and below, HP denotes hyperparameters and HPO denotes hyperparameter optimization. The shot count refers

Cohort or analysis

Users/models

Flickr-AES users Candidate FL / unseen evaluation users Eligible FL clients per round Global-validation monitoring identities 10-shot enhancers: completed / skipped 100-shot enhancers: completed / skipped Image-suitability analysis

210 173 / 37 87 10 36 / 1 37 / 0 33

Table 6: Cohort accounting after the stated eligibility controls. Each analysis uses a fixed matched cohort for all reported comparisons. only to rated support images used for gradient-based scorer adaptation. Validation data are used for early stopping or model selection, and the test split is disjoint from both. A personalized scorer is eligible to guide the enhancer only if its validation Spearman rank correlation (SRCC) is at least 0.10, ensuring a uniform quality threshold for enhancer guidance. The resulting cohort counts are reported in Tab. 6, and each analysis keeps its eligible cohort fixed across all compared configurations.

E.3

Training and Model Selection

Randomness is controlled by seeding Python’s random, NumPy, the PyTorch CPU generator, and all CUDA generators. Dataset and cohort splitting, as well as the federated and centralized scorer runs, use seed 42. Per-user scorer calibration uses seed 42+u for numeric client identifier u. Fixed-HP enhancer adaptation, shared enhancer HPO, and enhancement evaluation use seed 60. Unless otherwise stated, each internally trained model configuration is trained once under these fixed seeds. Personalized aggregate rows contain one trained model per eligible user. The HPO trial counts reported below are search trials rather than repeated training seeds. Global Models. The same pretrained CLUT-Net checkpoint is used unchanged as ϕg for all downstream experiments. The associated Paired Training Objective uses the unclipped output and combines L1 with 0.1Lperc , where LPIPS denotes learned perceptual image patch similarity. The global scorer is trained for 20 federated rounds with all eligible clients participating. Each client uses adaptive local epochs. The returned aggregation count mk is therefore the number of examples actually processed across its local optimization steps, as made explicit in Algorithm 1. Square-root weighting is applied to mk . Round 13 is selected by the highest global validation Spearman rank correlation (SRCC). Personalized Scorer Calibration. For each unseen user, a balanced support set of 10 or 100 local ratings is used for adaptation, and a separate validation split selects the checkpoint. Per-user scorer HPO uses 20 trials per user to maximize validation SRCC. MSE and Pearson linear correlation (PLCC) are reported metrics rather than components of the search objective. The search covers learning rate [10−5 , 5 × 10−4 ] on a log scale, pairwise weight [0, 0.30], weight decay [10−7 , 10−3 ], variance weight [0, 0.05], and gradient clipping [0.5, 2.0]. The regression coefficient is

Hyperparameter

10-shot

100-shot

LR λpref λaes λ1 λperc λgap µ Gradient clip

5.4485×10−4 0.0411 0.5996 0.1007 0.0543 0.5107 0.1048 1.9631

5.3948×10−4 0.0402 0.5848 0.1271 0.0543 0.7642 0.2051 1.6176

Table 7: Selected shared enhancer HPO configurations. Both use 40 final epochs. The value µ is the excess-gap tolerance. The search used 20 trials and selected trial 12 in both support regimes. max(0.60, 1 − λpair ). Epoch ranges are 20–40 for 10-shot and 30–80 for 100-shot, with early stopping. The final metrics are evaluated on the disjoint user test split. Frozen-Scorer-Guided Enhancer Adaptation. All four reported enhancer conditions use the corresponding personalized scorers obtained by per-user scorer HPO. The fixed-HP and shared enhancer HPO labels therefore refer only to Frozen-Scorer-Guided Enhancer Adaptation, not to different Personalized Scorer Calibration strategies. The fixed-HP gain-selection configuration operates at 224 × 224 with batch size 4, Adam learning rate 3 × 10−4 , and 40 epochs. It uses λaes = 0.5, λ1 = 0.1, λperc = 0.05, λgap = 3.0, excess-gap tolerance µ = 0.05, and gradient clipping 1.0. The preference coefficient λpref is chosen from {0.01, 0.03, 0.05, 0.09} according to validation-SRCC intervals [0.10, 0.20), [0.20, 0.30), [0.30, 0.40), and [0.40, 1], respectively. Scorers below 0.10 are skipped. Checkpoints maximize validation preference gain, corresponding to γ1 = γp = 0 in the main-paper selection criterion. For shared-HPO regularized selection, 20 enhancer HPO trials are evaluated on the validation partitions of eight users. Tab. 7 gives the selected settings, which are then used for 40epoch personalization. Checkpoints maximize the regularized validation criterion from the main paper using positive fidelity penalties. The eight validation identities are drawn from the final 37-user cohort, while all test images remain held out. The selected configuration is then fixed across the cohort for the within-cohort trade-off and matched objective analyses. The prespecified fixed-HP setting provides the user-disjoint open-world reference. Training, fine-tuning, and evaluation were conducted on a Windows 11 workstation equipped with an Intel Core Ultra 9 285K CPU, an NVIDIA RTX 5090 GPU, approximately 64 GB of system memory, and CUDA 12.8. The experimental environment uses PyTorch 2.2.2 and Torchvision 0.17.2. Preprocessing also used Apple M2 and M3 devices. The deployed enhancer contains 0.293M parameters, with a measured forward time of 4.62 ms per 224 × 224 image.

E.4

Resource Accounting

We derive the following resource counts from the implemented modules and update masks. Parameter totals are exact. Arithmetic counts are static estimates that use one

Stage

Resident Updated FLOPs/sample

Federated Aesthetic Preference Learning Personalized Scorer Calibration (10-shot) Personalized Scorer Calibration (100-shot) Frozen-Scorer-Guided Enhancer Adaptation Personalized On-Device Inference

0.787M 0.787M

4.70M

0.787M 0.527M

3.66M

0.787M 0.787M

4.70M

6.523M 0.265M

5.16G

0.293M

Table 8: Complete stage-level resource summary. Resident counts all required weights, whereas Updated counts gradient-updated weights. Scorer rows assume cached descriptors. The Frozen-Scorer-Guided Enhancer Adaptation row includes frozen supervision networks. Model block

Parameters

Scorer color projection Scorer semantic projection Scorer fusion MLP Scorer temperature Lightweight Dual-Cue Aesthetic Scorer total

13,824 246,528 526,849 1

CLUT CNN backbone CLUT coefficient head Compressed LUT bases CLUT-Net total

245,504 19,092 27,945 292,541

787,202

Table 9: Exact parameter decomposition of the two trainable FedPAIE models. multiply–accumulate (MAC) as two floating-point operations. They exclude data loading, normalization, nonlinear activations, elementwise HSV/Lab operations, loss reductions, optimizer bookkeeping, and validation. Memory is reported in MiB and assumes FP32. It is an analytical modelstate lower bound rather than measured peak device memory because activations, CUDA workspaces, dataloader buffers, and retained checkpoint copies depend on the runtime. The 10-shot scorer mask updates the fusion MLP and temperature, totaling 526,850 parameters. The 100shot mask updates all 787,202 scorer parameters. FrozenScorer-Guided Enhancer Adaptation freezes the 27,945 compressed-basis parameters and updates the 245,504parameter backbone and 19,092-parameter coefficient head, totaling 264,596 parameters. The complete Frozen-Scorer-Guided Enhancer Adaptation stack in Tab. 10 consists of the 0.293M-parameter CLUT-Net, 0.787M-parameter personalized scorer, 2.972Mparameter MobileNetV3 feature trunk, and 2.471Mparameter AlexNet-LPIPS network. The latter three modules and the compressed LUT bases are frozen. Only the 0.265Mparameter CLUT coefficient predictor is optimized. Representative serialized files occupy 3.092 MiB for the global scorer, 3.012 MiB for a personalized scorer, 1.129 MiB for the generic CLUT checkpoint, and 1.125 MiB for a personalized CLUT checkpoint. Small differences from raw FP32 weight size arise from serialization metadata.

Stage

Params. R/U

MiB W/T

Federated Aesthetic Preference Learning 787,202 / 787,202 3.003 / 12.012 Personalized Scorer Calibration (10-shot) 787,202 / 526,850 3.003 / 9.035 Personalized Scorer Calibration (100-shot) 787,202 / 787,202 3.003 / 12.012 Frozen-Scorer-Guided Enhancer Adaptation 6,522,543 / 264,596 24.882 / 27.910 Personalized On-Device Inference 292,541 / – 1.116 / 1.116

Table 10: Parameter-state memory across the three optimization stages and deployment. R/U denotes resident/updated parameters, and W/T denotes FP32 weights/total training state. The latter includes gradients and two optimizer moment tensors for updated parameters. Activations and workspaces are excluded. Operation

MACs / FLOPs

Scope

Scorer forward 0.783M / 1.565M All projections and fusion MLP Federated Aesthetic Prefer- 2.348M / 4.696M Forward and backward through ence Learning all scorer blocks Personalized Scorer Calibra- 1.832M / 3.663M Frozen projections and trainable tion (10-shot) fusion MLP CLUT image-transform path

0.179G / 0.358G Forward and gradients to the coefficient predictor Scorer-guidance path 0.431G / 0.862G MobileNetV3 and scorer gradient to the enhanced image LPIPS path 1.968G / 3.936G Two feature passes and enhanced-branch backward Frozen-Scorer-Guided En- 2.578G / 5.156G One 224 × 224 training image hancer Adaptation

Table 11: Static core-arithmetic accounting. Batch size changes parallelism but not the per-sample values. The estimate excludes the operations listed in the opening paragraph of this subsection. Federated Aesthetic Preference Learning. The scorer optimizer consumes cached 24-D HSV/Lab descriptors and 960-D MobileNetV3 embeddings. Each FP32 descriptor pair occupies approximately 3.84 KiB per image. If the semantic embedding is not already cached, its one-time local extraction activates a 2.972M-parameter MobileNetV3 trunk and requires approximately 0.215 GMAC per image. The resulting embedding is reused across every local epoch and communication round. For client k with Nk cached training samples, Ek local epochs, and T = 20 communication rounds, the optimization-only computation is: CkFL ≈ T Ek Nk (4.696 MFLOPs).

(34)

Eligible clients with 100 ≤ Nk < 500 use two local epochs, while larger clients use one. For example, Nk = 100 and Ek = 2 require approximately 0.94 GFLOPs per round and 18.78 GFLOPs over 20 rounds. The FP32 scorer state is 3.149 MB, or 3.003 MiB, per download or upload. Twenty rounds therefore transfer approximately 62.98 MB in each direction and 125.95 MB bidirectionally per participating

client. Personalized Scorer Calibration. For a support set of Su samples and Eu calibration epochs, the optimization-only computation is: Cucal ≈ Su Eu cu ,  cu =

3.663 MFLOPs, 4.696 MFLOPs,

Su = 10, Su = 100.

(35)

The 20–40 epoch range for 10-shot calibration corresponds to 0.73–1.47 GFLOPs for one final fit. The 30–80 epoch range for 100-shot calibration corresponds to 14.09– 37.56 GFLOPs. The reported per-user scorer HPO path runs 20 sequential trials followed by one final refit. Without early stopping, its optimization-only envelope is 15.39– 30.77 GFLOPs for 10-shot and 295.82–788.85 GFLOPs for 100-shot. Early stopping reduces the actual total. These envelopes exclude repeated validation inference. Sequential trials increase total arithmetic but do not multiply the peak model state in Tab. 10. Frozen-Scorer-Guided Enhancer Adaptation. The original-image scorer outputs are cached, but features of the changing enhanced image cannot be cached. We therefore count gradients through the frozen MobileNetV3 trunk and personalized scorer. LPIPS similarly performs feature extraction for the original and enhanced images, with backward propagation through the enhanced branch. For Nu unpaired local images and the reported 40-epoch schedule, the optimization-only computation is: CuE ≈ 40Nu (5.156 GFLOPs).

(36)

Local sets of 25, 50, and 100 images therefore require approximately 5.16, 10.31, and 20.62 TFLOPs for one final adaptation. The reported shared enhancer HPO uses 20 trials, 10 epochs per trial, and at most eight representative users.PIts corresponding optimization-only budget is 20 × 10 × u Nu × 5.156 GFLOPs and is separate from the final 40-epoch adaptations. These static counts describe arithmetic rather than wall-clock latency. They show that the trainable state remains compact even when all frozen supervision networks are conservatively included.

E.5

Metric and Reference Conventions

For target ratings y and predictions ŷ, scorer evaluation uses:  ρS = corr rank(y), rank(ŷ) , rP = corr(y, ŷ), (37) where corr denotes Pearson correlation. Here, ρS and rP are SRCC and PLCC, respectively. MSE is also reported. For user u, the enhancement proxy is: ∆u (I) = Sθu (Eϕu (I)) − Sθu (I).

(38)

A positive value means that the same frozen personalized scorer used for guidance assigns a higher score to the enhanced image. This optimization-aligned personalization proxy is complemented by reference-based fidelity metrics and paired user-level tests. The image-reference metrics are peak signal-to-noise ratio (PSNR), structural similarity (SSIM), and LPIPS.

Quantity

Interpretation

SRCC / PLCC / MSE

Agreement between predicted and observed user ratings. Personalized scorer proxy. Higher is better under that scorer.

Predicted score / scorerpredicted preference gain ∆ PSNR/SSIM/LPIPS vs. input PSNR/SSIM/LPIPS vs. Expert C Inter-client output difference

Content preservation and magnitude of the applied transformation. Similarity to a standardized professional retouching reference. Personalization diversity among user-specific outputs for a shared input.

Table 12: Evaluation quantities and the claims they support. Metric

Initial

Reported final/peak

Training loss Validation MSE Validation SRCC Validation PLCC Prediction standard deviation

0.0694 0.0808 0.4650 0.4764 0.0861

0.0401 (round 20) 0.0603 (round 18 minimum) 0.5623 (round 13 peak) 0.5854 (round 20 maximum) 0.1769 (0.1771 peak)

Table 13: Global scorer behavior during Federated Aesthetic Preference Learning over 20 rounds. Unless otherwise stated, reported personalized results are means over the valid client cohort. PSNR and SSIM are better when larger, whereas LPIPS is better when smaller. Their meaning depends on the explicitly named reference.

F.1

F

Extended Quantitative Results

Federated Aesthetic Preference Learning and Personalized Scorer Calibration

Tab. 13 summarizes the global trajectory. Validation MSE stabilizes after approximately 10–12 rounds. The small lateround oscillations are consistent with non-IID data and stochastic local minibatches. Because all eligible clients participate in every round, the trajectory reflects optimization under the full eligible federated cohort. With 10 support ratings, the federated initialization with fixed calibration hyperparameters gives higher SRCC and PLCC than the centralized counterpart, while the centralized model has lower MSE. With 100 ratings, the centralized model is marginally strongest on all three metrics. Both Support

Scorer initialization

SRCC↑

PLCC↑

MSE↓

10 10 10

Centralized Federated, fixed HP Federated, per-user HPO

0.5291 0.5412 0.5413

0.5404 0.5471 0.5375

0.0532 0.0589 0.0700

100 100 100

Centralized Federated, fixed HP Federated, per-user HPO

0.5690 0.5649 0.5625

0.5727 0.5665 0.5646

0.0552 0.0595 0.0609

Table 14: Mean test results after Personalized Scorer Calibration on unseen users.

CLUT training target Expert A Expert B Expert C Expert D Expert E Mixed-style

PSNR (dB)↑ 22.28 27.36 25.21 24.06 25.31 23.69

Table 15: Training-target comparison for expert-specific and mixed-style CLUT models under their respective retouch targets. federated variants improve their correlations when the support set increases from 10 to 100 ratings. Fixed-HP calibration remains competitive with per-user scorer HPO in both regimes. For the federated initialization with fixed calibration hyperparameters, the user-level SRCC standard deviations are 0.130 and 0.116 in the 10- and 100-shot settings, respectively.

F.2

Enhancer Adaptation Strategy Comparison

Tab. 16 compares three adaptation and checkpoint-selection strategies: absolute-score selection, fixed-HP gain selection, and shared-HPO regularized selection. Input-reference metrics measure preservation rather than absolute enhancement quality. For the first two strategies, the displayed score and fidelity values come from test data, whereas the bracketed preference gains are validation quantities used during selection. The shared-HPO strategy reports both proxy quantities on held-out test images. The superscripts make each quantity’s evaluation role explicit, and the analysis compares checkpoint-selection behavior across the three strategies. Absolute-score selection yields smaller validation preference gains and lower input fidelity than fixed-HP gain selection in both support regimes, indicating that absolute score alone is a weaker selection signal for user-specific improvement. Fixed-HP gain selection is the strict open-world reference: it preserves the input closely and produces a positive validation preference gain with both support sizes. SharedHPO regularized selection demonstrates stronger optimization of the scorer proxy and permits a stronger transformation, occupying a more preference-oriented operating point. With 100 ratings, fixed-HP gain selection improves inputreference PSNR from 30.43 to 31.12 dB, increases SSIM from 0.9627 to 0.9716, and reduces LPIPS from 0.0162 to 0.0132. Shared-HPO regularized selection raises the mean test score from 0.5624 to 0.5864 and the mean test preference gain from 0.0826 to 0.0858. These aggregate changes support the value of additional ratings while preserving the

Score [∆]↑

P/S/L-in

0.5483t [0.0083v ] 0.5039t [0.0247v ] 0.5624t [0.0826t ]

27.30/0.9515/0.0308 30.43/0.9627/0.0162 20.51/0.7665/0.1206

10 10 10

Absolute-score selection Fixed-HP gain selection Shared HPO + reg. selection

100 100 100

Absolute-score selection 0.5519t [−0.0004v ] 29.18/0.9643/0.0234 Fixed-HP gain selection 0.5245t [0.0244v ] 31.12/0.9716/0.0132 Shared HPO + reg. selection 0.5864t [0.0858t ] 19.76/0.7419/0.1348

Table 16: Comparison of enhancer-adaptation and checkpoint-selection strategies. Superscripts v and t denote validation- and test-split proxy values. Fidelity columns use the input as reference, and P/S/L denotes PSNR/SSIM/LPIPS. Bracketed values are preference gains, not standard deviations. The shared-HPO configuration is selected on validation partitions and evaluated on held-out test images.

Generic Enhancement Prior Learning

Tab. 15 reports a training-target comparison for five expertspecific CLUT models and one mixed-style model. Each expert-specific model is evaluated against that expert’s retouches, while the mixed-style configuration uses its associated retouch targets. Because the target distributions differ across rows, these results characterize target-specific reconstruction behavior rather than a common-reference ranking.

F.3

Support Strategy

Support

Enhancer

PSNR-C↑

SSIM-C↑

LPIPS-C↓

10 10 100 100

Fixed HP HPO Fixed HP HPO

17.97 18.33 18.12 18.58

0.801 0.777 0.804 0.775

0.130 0.147 0.129 0.147

Table 17: FedPAIE evaluated against Expert C retouches, which provide a standardized generic professional reference. distinction between fixed-HP fidelity and shared-HPO proxypreference strength.

F.4

Detailed Enhancement Baselines and Protocol Context

The main paper presents the enhancement comparison in a compact consolidated table. Here we separate results obtained under the common evaluation protocol from literaturereported operating points so that the source of every number remains explicit. Tab. 18 is the controlled comparison: all rows use the same Expert C reference and evaluation pipeline. Relative to personalized AdaInt, 10- and 100-shot FedPAIE increase SSIM by 0.045 and 0.048 (approximately 6.0% and 6.3%) and reduce LPIPS by 0.101 and 0.102 (approximately 43.7% and 44.2%). AdaInt retains PSNR advantages of 1.48 and 1.33 dB. Because all three metrics use Expert C as a generic professional reference, these differences characterize a fidelity trade-off rather than direct evidence of Method Original input AdaInt + personalized scorer (Yang et al. 2022a) FedPAIE, 10-shot fixed HP FedPAIE, 100-shot fixed HP

PSNR-C↑ SSIM-C↑ LPIPS-C↓ 17.84 19.45 17.97 18.12

0.791 0.756 0.801 0.804

0.138 0.231 0.130 0.129

Table 18: Controlled Expert C comparison under the common evaluation protocol. AdaInt has the highest PSNR, while 100-shot FedPAIE has the highest SSIM and lowest LPIPS.

Literature-reported method

PSNR

SSIM

Parameters

SpliNet (Bianco et al. 2020) PieNet (Kim, Koh, and Kim 2020) Masked Style Modeling (Kosugi and Yamasaki 2024)

18.74 20.52 22.98

0.819 0.850 0.897

0.03M 28M 90M

Table 19: Literature-reported personalized enhancement results under their original evaluation protocols. The table provides broader quality and model-scale context for the controlled comparison in Tab. 18. user-preference superiority. Tab. 19 complements the controlled comparison with representative personalized enhancement results reported in prior work. The rows retain their original datasets, targets, resolutions, and evaluation implementations. They therefore characterize the broader quality–efficiency landscape, while Tab. 18 provides the direct method comparison. We omit unreported or speculative latency estimates. For SpliNet, PSNR and SSIM follow the 20-preference FiveK evaluation reported by Kosugi and Yamasaki (Kosugi and Yamasaki 2024). Its 0.03M parameter count is computed from the official 10-node, eight-base-filter personalized architecture (Bianco et al. 2020), which contains approximately 31.4K trainable parameters, rather than quoted from the original paper. SpliNet is therefore smaller than FedPAIE in raw trainable-parameter count. The two models have different transformation and system scopes: SpliNet predicts global per-channel neural-spline color transforms, whereas FedPAIE uses an image-adaptive 3D-LUT enhancer in a pipeline that connects Federated Aesthetic Preference Learning with local user adaptation. The 0.293M FedPAIE figure thus describes a lightweight image-adaptive enhancer in a broader federated personalization pipeline, while SpliNet remains the smallest architecture in Tab. 19.

G G.1

Ablation, Qualitative, and Image-Suitability Analyses

Objective Ablation for Frozen-Scorer-Guided Enhancer Adaptation

This ablation acts only on Frozen-Scorer-Guided Enhancer Adaptation. It does not change Federated Aesthetic Preference Learning or Personalized Scorer Calibration. The 10- and 100-shot studies contain 36 and 37 eligible unseen users, respectively. Within each support regime, all variants use the same FiveK evaluation pairs, frozen personalized scorers, shared enhancer HPO configuration selected for the Full objective, validation-SRCC eligibility rule, hashverified original-score cache, training schedule, preprocessing, and checkpoint-resume policy. Eligibility is determined only from validation SRCC. Held-out test SRCC does not affect client inclusion. No removal variant re-runs shared enhancer HPO, and every retained coefficient is copied from Tab. 7. Fixing the shared configuration across all variants yields a matched within-cohort intervention in which the sole change is which terms of the Personalized Enhancement Objective remain active: The evaluation includes two reusable controls that require

Variant

Lpref

Laes

L1

Lperc

Lgap

Full objective Without reg. group Without scorer guidance Without excess-gap penalty Without Lpref

Yes Yes – Yes –

Yes Yes – Yes Yes

Yes – Yes Yes Yes

Yes – Yes Yes Yes

Yes – – – Yes

Table 20: Active objective terms for Frozen-Scorer-Guided Enhancer Adaptation. “Without reg. group” removes the fidelity-plus-gap regularization group (L1 , Lperc , Lgap ). no additional training. Original is the unenhanced input, and Generic Enhancement Prior denotes the shared pretrained CLUT-Net checkpoint obtained through Generic Enhancement Prior Learning without personalization. The Full objective and four removal variants use 500 common MIT-Adobe FiveK images with Expert C input–target pairs for every eligible user. Each user’s frozen personalized scorer supplies the proxy axis. PSNR, SSIM, and LPIPS relative to Expert C supply the fidelity axis. Reporting both axes jointly characterizes preference improvement and image preservation at each operating point. Personalization Relative to the Controls. Tab. 21 and Fig. 6 show that the Full objective improves the frozen personalized scorer output over Original by 0.0890 and 0.0904 for 10 and 100 support ratings. It also improves over the Generic Enhancement Prior by 0.0274 and 0.0285. Tab. 22 confirms these differences with paired user-level tests. The Full objective exceeds the Generic Enhancement Prior for every evaluated user, namely 36 of 36 in the 10-shot setting and 37 of 37 in the 100-shot setting. This universal direction under the frozen personalized scorer proxy provides stronger user-level evidence than an aggregate mean alone. Together with the fidelity axes and qualitative views, it supports the same preference–fidelity operating point. Fig. 11 provides two additional 100-shot examples. Scorer Guidance Supplies the Personalization Signal. Without scorer guidance, only L1 + Lperc remain. The mean gain becomes 0.0002 in the 10-shot setting and 0.0001 in the 100-shot setting. On Flickr-AES, the same variant reaches 74.1 and 73.6 dB PSNR relative to the input with SSIM of approximately 0.9999. It therefore converges to an identitylike transformation. The 81% and 68% positive-direction rates in Tab. 21 correspond to minute changes around zero, not meaningful personalization. This variant provides a direct lower bound and isolates the frozen personalized scorer as the source of measurable proxy-guided adaptation. The Excess-Gap Penalty Protects Fidelity. Removing only Lgap raises the proxy score from 0.5213 to 0.5295 for 10shot and from 0.5438 to 0.5513 for 100-shot. Read alone, these numbers would incorrectly favor the removal variant. The fidelity axis reveals the failure mode. Relative to the Full objective, PSNR falls by 4.90 and 4.35 dB, SSIM falls by 0.282 and 0.241, and LPIPS increases by 0.194 and 0.163. This controlled result gives Lgap a clear component-level interpretation. It limits gains that exploit the scorer proxy and preserves a usable fidelity floor. The Fidelity-Plus-Gap Regularization Group Improves

10-shot (n = 36) Variant Original Generic prior Full objective w/o Lpref w/o Lgap w/o reg. group w/o scorer guidance

100-shot (n = 37)

Score [∆]

Win

PC

SC

LC

Score [∆]

Win

PC

SC

LC

0.4324 [0.0000] 0.4939 [+0.0615] 0.5213 [+0.0890] 0.5160 [+0.0836] 0.5295 [+0.0971] 0.5173 [+0.0850] 0.4326 [+0.0002]

– 100 100 100 100 94 81

– 22.60 18.33 18.46 13.43 7.30 18.05

– 0.904 0.777 0.789 0.495 0.235 0.814

– 0.087 0.147 0.139 0.341 0.548 0.122

0.4534 [0.0000] 0.5153 [+0.0619] 0.5438 [+0.0904] 0.5436 [+0.0903] 0.5513 [+0.0979] 0.5269 [+0.0736] 0.4535 [+0.0001]

– 100 100 100 100 86 68

– 22.60 18.58 18.45 14.23 7.73 18.04

– 0.904 0.775 0.775 0.534 0.289 0.814

– 0.087 0.147 0.145 0.310 0.529 0.122

Table 21: Complete dual-axis objective ablation on 500 FiveK images. Score is the mean frozen personalized scorer output, and ∆ is the scorer-predicted preference gain from Original, and the win rate is the fraction of users with a positive mean change. Personalized rows aggregate 36 × 500 or 37 × 500 outputs. The Generic Enhancement Prior fidelity metrics are computed once over the 500 images and shared across users. Expert C provides a standardized professional reference. Displayed statistics are rounded independently from unrounded aggregates. Bracketed values are gains, not standard deviations. No Gap

0.100

No Gap

0.100

Full

No Reg

Personalized aesthetic-score gain

Personalized aesthetic-score gain

Full

No Rank

0.075

Global 0.050

0.025

0.000

No Rank 0.075

No Reg Global

0.050

0.025

0.000 8

12

No Scorer 20

16

24

PSNR to FiveK Expert-C (dB)

8

12

16

No Scorer 20

24

PSNR to FiveK Expert-C (dB)

Figure 6: Scorer-predicted preference-gain–fidelity trade-off in the objective ablation for Frozen-Scorer-Guided Enhancer Adaptation on 500 FiveK images. The left and right panels show the 10-shot and 100-shot settings, respectively. Each labeled point relates mean gain from Original under the same frozen personalized scorer to PSNR relative to Expert C. Global denotes the Generic Enhancement Prior, No Rank denotes w/o Lpref , No Gap denotes w/o Lgap , No Reg denotes w/o the fidelityplus-gap regularization group, and No Scorer denotes w/o scorer guidance. Full improves the scorer proxy beyond the Generic Enhancement Prior while retaining substantially higher fidelity than No Gap and No Reg. The score is an optimization-aligned user-specific proxy, and Expert C provides a standardized professional reference.

Support Comparison 10 10 10 100 100 100

Prior–Original Full–Original Full–Prior Prior–Original Full–Original Full–Prior

Mean diff.

t

p

0.0615 0.0890 0.0274 0.0619 0.0904 0.0285

19.01 12.43 5.49 26.65 19.37 7.97

−20

5.0 × 10 2.1 × 10−14 3.6 × 10−6 2.7 × 10−25 1.3 × 10−20 1.8 × 10−9

Table 22: Paired user-level tests on frozen personalized scorer outputs. The tests quantify consistency across users under the learned personalization proxy. Prior denotes the Generic Enhancement Prior. The support cohorts contain 36 and 37 paired users.

Cross-Distribution Behavior. Jointly removing L1 , Lperc , and Lgap produces the highest in-domain Flickr-AES scores, 0.6923 and 0.7019, but simultaneously reduces inputreference PSNR to 7.42 and 7.95 dB. On FiveK, its scores fall below the Full objective while its PSNR relative to Expert C reaches only 7.30 and 7.73 dB. The apparent in-domain advantage therefore does not transfer. Full improves every reported proxy and fidelity measure over this removal variant. It restores 11.03 and 10.85 dB PSNR and reduces LPIPS by 0.401 and 0.382. Tab. 23 shows the full comparison. These results support the fidelity-plus-gap regularization group as a joint defense against proxy exploitation. The comparison between the variant without Lgap and the variant without the complete fidelity-plus-gap regularization group further isolates the two fidelity losses as a pair. Retaining L1 + Lperc improves FiveK PSNR by 6.13 and 6.50 dB, increases SSIM by 0.260 and 0.245, and reduces LPIPS by 0.207 and 0.219. It also raises the FiveK proxy score by 0.0122 and 0.0244. Thus,

Variant Full objective w/o Lpref w/o Lgap w/o reg. group w/o scorer guidance

Flickr score / Pin 10-shot / 100-shot

FiveK score 10-shot / 100-shot

0.5624/20.51 0.5864/19.76 0.5575/21.10 0.5864/19.90 0.6448/13.30 0.6571/14.44 0.6923/7.42 0.7019/7.95 0.4794/74.1 0.5001/73.6

0.5213 0.5438 0.5160 0.5436 0.5295 0.5513 0.5173 0.5269 0.4326 0.4535

Table 23: Cross-distribution analysis. Flickr-AES values use held-out in-domain test images, and PSNR-in measures change from the input. FiveK scores use the 500-image evaluation set. Each stacked cell lists the 10-shot value above the 100-shot value. A high scorer output accompanied by very low input fidelity and weak transfer is consistent with proxy exploitation. the two fidelity losses jointly improve both the usable operating point and cross-distribution stability rather than merely suppressing enhancement. The paired intervention evaluates L1 and Lperc as the intended fidelity component, matching their joint role in preserving pixel and perceptual content. The Flickr-AES-to-FiveK score decrease is 0.0411 and 0.0426 for Full, compared with 0.1153 and 0.1058 without Lgap and 0.1750 in both regimes without the complete fidelity-plusgap regularization group. This gap provides an additional transfer-oriented view of the protective terms. Fig. 7 visualizes the same controlled cross-distribution comparison. The Preference Terms Are Compatible. Removing Lpref reduces the 10-shot score from 0.5213 to 0.5160, while the 100-shot score changes from 0.5438 to 0.5436. Fidelity stays close in both regimes. At the selected weights, Laes already provides a strong absolute preference signal, while Lpref explicitly encodes the desired ordering between enhanced and original images. Its additional effect is most visible under sparse support, while preserving the 100-shot solution. The two preference-bearing terms therefore provide overlapping and compatible supervision. Why the Generic Enhancement Prior Has Higher Fidelity to Expert C. The Generic Enhancement Prior is initialized from paired FiveK professional retouches, whereas personalized adaptation deliberately moves its output toward each user’s frozen personalized scorer. The Expert C axis is therefore naturally favorable to the generic professional prior. Its PSNR of 22.60 dB should be read together with the Full objective’s higher personalized score and consistent per-user advantage. The result characterizes the intended personalization–reference trade-off.

G.2

Orthogonal Optimization-Choice Ablation

The loss-component ablation above fixes the optimization configuration and changes only the active objective terms. We separately vary enhancer hyperparameters, Personalized Scorer Calibration, and support size. These factors form a 2× 2×2 design and answer a different question from component attribution.

Enhancer HP

Scorer calibration 10-shot score / PSNR-in 100-shot score / PSNR-in

Fixed HP

Fixed HP Per-user scorer HPO

Fixed HP Shared enhancer HPO Fixed HP Shared enhancer Per-user scorer HPO HPO

0.499 / 28.4

0.510 / 30.8

0.504 / 30.4

0.525 / 31.1

0.547 / 18.5

0.572 / 19.3

0.562 / 20.5

0.586 / 19.8

Table 24: Orthogonal 2×2×2 optimization-choice ablation. Scores and PSNR-in are measured on the Flickr-AES test split, with the input as the PSNR reference. HP and HPO denote hyperparameters and hyperparameter optimization. Tab. 24 shows that per-user scorer HPO in Personalized Scorer Calibration improves the predicted score under both enhancer configurations and both support sizes. It also improves input-reference PSNR in every matched comparison. Shared enhancer HPO produces the larger score increase and intentionally permits a stronger transformation, which lowers PSNR-in relative to the fixed-HP configuration. Combining the two choices gives the highest frozen personalized scorer output at both support sizes while recovering some fidelity relative to shared enhancer HPO with fixed-HP scorer calibration. This factorial result supports the selected configuration without conflating optimization choices with losscomponent evidence. Results for Federated Aesthetic Preference Learning and Personalized Scorer Calibration remain in Tab. 14, where the federated initialization with fixed-HP calibration reaches SRCC 0.541±0.130 and 0.565±0.116 for 10 and 100 support ratings. Together, these experiments evaluate the complete Lightweight Dual-Cue Aesthetic Scorer under both sparse and richer support.

G.3

Qualitative and Client-Level Analyses

The main paper presents the complete shared-input comparison. Across Examples A–E, Clients 104, 108, and 172 favor cooler outputs, Client 123 stays closer to the Generic Enhancement Prior, and Client 153 applies a darker, highercontrast transform. Fig. 9 adds a separate user-level view by placing rating-5 Flickr-AES examples beside unpaired FiveK inputs and Full-objective outputs. Fig. 8 shows that tonal and chromatic differences remain visible in local crops while scene structure is preserved. Fig. 10 extends the analysis beyond the displayed clients. The nonuniform pairwise distances show that output diversity is distributed across the full 100-shot evaluation cohort rather than confined to the selected qualitative examples. These figures jointly document distinct transformations and user-specific preference context. To complement the main-paper user 204 example, the additional visualizations use users 199 and 26. On the displayed images, Full raises the frozen-scorer output from 0.535 to 0.626 and from 0.339 to 0.472 relative to the Generic Enhancement Prior, while PSNR/SSIM increase from 31.2/0.94 to 31.6/0.95 and from 23.4/0.88 to 26.3/0.92, respectively. Fig. 11 shows both cases, and Fig. 9 provides complementary preference context. Fig. 12 extends the shared-input comparison with a randomly selected sheet

Mean scorer output on FiveK images

10-shot (𝑛 = 36)

100-shot (𝑛 = 37)

0.7

0.7

0.65

0.65

0.6

0.6

0.55

0.55

0.5

0.5

0.45

0.45 0.45

0.5

0.55

0.6

0.65

0.7

Mean scorer output on held-out Flickr-AES images

0.45

0.5

0.55

0.6

0.7

0.65

Mean scorer output on held-out Flickr-AES images

Full objective Without Lpref Without Lgap Without regularization group Without scorer guidance The dashed line marks equal mean scorer outputs on the two image distributions.

Figure 7: Cross-distribution analysis for the objective ablation of Frozen-Scorer-Guided Enhancer Adaptation. Each point compares the mean frozen personalized scorer output on held-out Flickr-AES images with that on the common 500-image FiveK set. The dashed line marks equal mean outputs on the two image distributions. The Full objective remains nearer this line than the variants without Lgap or the complete fidelity-plus-gap regularization group. Their larger decreases, together with the low input-reference fidelity in Tab. 23, are consistent with non-transferable proxy over-optimization. Together with the paired tests and fidelity metrics, this analysis provides complementary evidence across image distributions. containing five clients. Python’s random.Random(42) selects one of three pre-generated, non-overlapping candidate cohorts, yielding Clients 194, 199, 200, 204, and 210. Aggregate claims rely on the complete-cohort tables and matrix.

G.4

Image-Conditioned Output Variation

For a common input I and U personalized enhancers, we define the average pairwise output difference as: X 2 D(I) = MAE(Eϕu (I), Eϕv (I)) . (39) U (U − 1) u<v Here u and v index personalized users and MAE is the mean absolute RGB error between two outputs. The analysis applies 33 valid 10-shot enhancers to 500 inputs. D(I) ranges from 0.011 to 0.073. The 15 lowest- and 15 highest-difference inputs form the LOW and HIGH groups, respectively. Tab. 25 shows that the HIGH group is brighter, less saturated, and lower contrast on average. These statistics characterize the joint attribute profile of the two groups. Fig. 13 complements this input-level grouping by summarizing how each displayed output changes brightness, contrast, saturation, and colorfulness relative to the input. We additionally compare proxy scores for the LOW and HIGH groups (Tab. 26). FedPAIE produces the largest HIGH–LOW change, 0.0904, exceeding SpliNet by 31.2% and PieNet by 10.6%. This result supports the central observation that FedPAIE expresses a larger image-conditioned proxy response when an input offers more editable color

Group LOW D(I) HIGH D(I)

Brightness

Saturation

Contrast

0.167 0.346

0.402 0.195

0.145 0.111

Table 25: Mean input attributes for the lowest- and highestdifference groups. Method

LOW

HIGH

HIGH−LOW

FedPAIE SpliNet (Bianco et al. 2020) PieNet (Kim, Koh, and Kim 2020)

0.4404 0.4990 0.4983

0.5308 0.5679 0.5800

0.0904 0.0689 0.0817

Table 26: Reported proxy-score response across image groups. FedPAIE shows the largest HIGH–LOW change in this analysis. space. Because identical scorer-calibration and generation protocols are not established for all three methods, the absolute score levels retain their method-specific context. The relative response between LOW and HIGH groups is the informative comparison here.

H

Evaluation Scope and Evidence Interpretation

The experimental suite combines a user-disjoint open-world reference, matched objective interventions, optimizationaligned personalization measures, and reference-based fi-

Global

Client 104

Client 108

Client 123

Client 153

Client 172

Detail

Example C

Detail

Example B

Detail

Example A

Input

Figure 8: Full-image and detail-crop comparison for Examples A–C. Yellow boxes mark the enlarged regions. The crops make client-dependent tonal and chromatic shifts visible around the wheel, clothing texture, and building facade while showing that local structures remain aligned. delity measures. The following scope clarifies how these complementary forms of evidence support the conclusions. Model Selection. Per-user scorer HPO uses support images for fitting and a disjoint validation split for selection, followed by evaluation on held-out images. Shared enhancer HPO selects one configuration from validation partitions of eight identities, with all test images held out. This configuration is fixed for within-cohort trade-off and objective analyses, while the prespecified fixed-HP results provide the user-disjoint open-world reference. The two tracks separate open-world generalization from controlled configuration analysis. Personalization Evidence. The frozen personalized scorer supplies both the training signal and the reported scorerpredicted preference gain, making ∆ a direct, optimizationaligned user-preference endpoint. Paired user-level tests quantify the consistency of this gain, while input- and ExpertC-reference metrics, cross-distribution analysis, and qualitative comparisons provide complementary evidence for fidelity and transformation stability. Blinded user evaluation offers a natural extension with direct perceptual feedback. The five-setting objective ablation for Frozen-Scorer-Guided Enhancer Adaptation is complete for the 10- and 100-shot settings and supports term-level conclusions for Lpref and Lgap and group-level conclusions for scorer guidance and fidelity-plus-gap regularization within its matched shared enhancer HPO analysis. The paired fidelity-loss intervention evaluates L1 and Lperc together as their functional preservation group, while the complete dual-cue scorer is evaluated throughout both support regimes. References and Matched Cohorts. Input-reference PSNR,

SSIM, and LPIPS measure preservation. Metrics relative to Expert C measure proximity to a standardized professional style, while the frozen personalized scorer provides the user-specific endpoint. The experimental design uses fixed, analysis-specific cohorts. The 10-shot evaluation uses 36 users, the 100-shot evaluation uses 37 users, and the imagesuitability analysis uses 33 valid enhancers. Within each objective analysis, identical eligibility rules and evaluation data are applied across all variants. This yields matched withinregime comparisons and descriptive trends across support scales. Comparison Context and Statistical Evidence. The literature rows position FedPAIE among representative recent personalized enhancement methods under their original protocols, while the study-evaluated rows provide same-protocol evidence. Paired tests cover the complete matched user cohorts, and the resource analysis reports parameters, trainable state, and static GFLOPs throughout the pipeline. Together, these results evaluate FedPAIE at the model, user, and image levels. Privacy Scope. Raw images and ratings remain local throughout the implemented protocol. Only scorer model updates are exchanged during federated training, establishing a clear protocol-level raw-data privacy boundary. Secure aggregation and differential privacy are compatible complementary communication-layer protections.

High Rate Image

Score 5

Score 5

Score 5

Enhanced Image

Original Image

Score 5

I.2

Figure 9: Additional preference context and Full-objective outputs for Flickr-AES user 26 after 100-shot Personalized Scorer Calibration. The top row shows four rating-5 Flickr-AES examples, the middle row shows four unpaired FiveK inputs, and the bottom row shows their outputs after Frozen-Scorer-Guided Enhancer Adaptation. The rating examples provide user-level preference context rather than paired target-style supervision.

I

Extended Related Work and Novelty Positioning

The main paper provides a compact account of the most relevant literature. This section expands that discussion by separating five dimensions that are often conflated: the aesthetic target, the form of preference supervision, the location of user data, the supervision used to learn an image transformation, and the model retained for inference. This separation is important because a method can be personalized without being federated, federated without producing an image, or computationally efficient without learning an individual user’s preference.

I.1

et al. 2022a), while CLUT-Net factorizes and compresses the LUT representation (Zhang et al. 2022). SVDLUT decomposes spatial-aware lookup tables to reduce model size and runtime while retaining spatial information (Kim, Lee, and Cho 2025). These LUT methods focus on efficient generic enhancement. FedPAIE instead uses a compressed representation as a preference-driven personalization engine within a federated pipeline. Generic paired retouches train the shared initialization. For a new user, the compressed LUT bases remain fixed while the lightweight coefficient predictor is adapted using a preference signal learned from local ratings. This design turns sparse private ratings into user-specific, image-adaptive 3D-LUT transformations without requesting paired retouches from the user.

Efficient Image Enhancement and Color Grading

Learning-based photo enhancement commonly estimates color and tone transformations from input–retouch pairs. MIT-Adobe FiveK established a standard paired setting with multiple expert renditions of each photograph (Bychkovsky et al. 2011). Subsequent systems increased content adaptivity through bilateral-grid prediction (Gharbi et al. 2017), removed strict pairing through adversarial learning (Chen et al. 2018), or incorporated an aesthetic objective into the enhancement process (Deng, Loy, and Tang 2018). These methods made learned enhancement more flexible, but their target is a generic enhancement distribution or an expert style. They do not infer the preference of an unseen user from that user’s sparse private ratings. Image-adaptive 3D lookup tables provide an especially efficient form of global color grading. Zeng et al. learn inputdependent mixtures of basis LUTs (Zeng et al. 2022). AdaInt learns nonuniform sampling intervals in color space (Yang

Personalized Aesthetics Assessment and Preference Learning

Generic aesthetic assessors estimate population-level quality or preference. NIMA predicts an aesthetic rating distribution (Talebi and Milanfar 2018), and VILA uses vision– language pretraining to learn aesthetics from user comments (Ke et al. 2023). Personalized image-aesthetics assessment instead models the deviation of an individual’s taste from a population prior (Ren et al. 2017). Representative approaches use collaborative attention (Wang, Yan, and Qin 2018), rich user and image attributes (Yang et al. 2022b), few-shot meta-learning (Zhu et al. 2022), graph-based collaboration (Shi et al. 2024), or multi-level transitional contrast learning that exploits cross-user contrastive information (Yang et al. 2024). Task-vector customization combines reusable vectors from generic-aesthetic and image-quality databases for scalable few-shot personalization and crossdomain generalization (Yun and Choo 2024). PAA+ extends the usual pretraining–fine-tuning paradigm with continual learning and validates it in a physique-aesthetics setting (Zhong et al. 2025). Recent datasets also broaden the assessment domain. LAPIS provides 11,723 artwork images with aesthetic ratings and rich image and personal attributes for personalized assessment (Maerten et al. 2025). Human preference comparisons provide a general mechanism for learning from limited feedback (Christiano et al. 2017), while user-guided reinforcement learning has been applied specifically to personalized image-aesthetics assessment (Lv et al. 2023). Together with the assessment methods above, these works show that personalized preferences can be learned from limited user feedback. The assessment methods themselves primarily output scores, rankings, or preference representations rather than enhanced images. Task-vector customization is especially close to our fewshot calibration setting, but it builds personalized scorers by centrally combining database task vectors. FedPAIE instead learns the shared scorer from decentralized ratings and uses aesthetic assessment as an intermediate model rather than the terminal task. Its lightweight dual-cue scorer combines color statistics with semantic context, which makes the prediction sensitive to both the appearance change and the image content. The global scorer is learned from decentralized image– rating pairs, after which the Support-Dependent Scorer Mask controls few-shot calibration for an unseen user. Pairwise or-

2 5 11 13 26 31 38 39

0.14

41 42 45

0.12

57 66

78

0.10

84

Client

86 93

0.08

99 104 108 113

0.06

123

Mean Absolute RGB Difference

75

127 131

0.04

132 142 153 170

0.02

172 179 181

0.00

194 199 200 204

210

204

200

199

194

181

179

172

170

153

142

132

131

127

123

113

108

99

104

93

86

84

78

75

66

57

45

42

41

39

38

31

26

13

5

11

2

210

Client

Figure 10: Pairwise output diversity over the full 100-shot evaluation cohort. Each cell is the mean absolute RGB difference between two personalized enhancers, averaged over the common FiveK inputs. The diagonal is zero by definition. Larger values indicate stronger client-dependent variation, while the scorer and fidelity analyses evaluate preference alignment and quality. dering and variance preservation are introduced only during this local calibration. The calibrated scorer is then frozen and used as a differentiable interface between sparse ratings and an image transformation. This role is distinct from reporting a personalized score. Gradients pass through the frozen personalized scorer to update the enhancer, so the preference estimate becomes a local transformation signal without making the scorer itself drift toward the images it rewards.

I.3

Personalized Image Enhancement

Personalized enhancement directly aims to produce different outputs for different users. Early work learns an individual preference model from user choices over candidate adjustments (Kang, Kapoor, and Lischinski 2010). PieNet represents a user’s taste as a preference vector and conditions a deep enhancement network on this representation (Kim, Koh, and Kim 2020). SpliNet embeds reference retouchers in a user space and predicts global neural-spline color transforms (Bianco et al. 2020). Masked style modeling further makes a user’s preferred style content-aware and trains with synthetically constructed input–retouch pairs (Kosugi and Yamasaki 2024). A recent Transformer approach infers

a preference vector from a few user-selected images and conditions enhancement on global and local style information (Kim et al. 2025). These methods demonstrate that a single generic output cannot represent the diversity of individual taste. Most closely, Personalized Photographic Style (PPS) learning infers a user’s photographic style from pairwise judgments and evaluates adaptations of style-transfer and enhancement models (Kim, Yoo, and Kim 2026). PerTouch uses semantic parameter maps and a VLM-driven agent with feedback-driven rethinking and scene-aware memory to translate language instructions and feedback into finegrained retouching controls (Chang et al. 2026). These methods establish nearby preference-to-transformation settings. PPS uses comparative judgments, while PerTouch uses language instructions and interactive feedback. Neither published formulation targets federated raw-data-local scorer learning or unpaired on-device enhancer adaptation. The supervision and privacy setting of FedPAIE are different. Prior personalized enhancers are generally trained in a centralized setting from explicit user choices, preferred-style examples, paired retouches, or pseudo-paired transforma-

(a)

(b)

(c)

(d)

(a)

(b)

(c)

(d)

Score 0.403

Reference

Score 0.535 PSNR 31.2 | SSIM 0.94

Score 0.626 PSNR 31.6 | SSIM 0.95

Score 0.185

Reference

Score 0.339 PSNR 23.4 | SSIM 0.88

Score 0.472 PSNR 26.3 | SSIM 0.92

(e)

(f)

(g)

(h)

(e)

(f)

(g)

(h)

Score 0.566 PSNR 30.3 | SSIM 0.91

Score 0.607 PSNR 26.2 | SSIM 0.88

Score 0.543 PSNR 8.7 | SSIM 0.56

Score 0.401 PSNR 21.4 | SSIM 0.84

Score 0.404 PSNR 21.1 | SSIM 0.87

Score 0.350 PSNR 14.9 | SSIM 0.77

Score 0.361 PSNR 4.2 | SSIM 0.19

Score 0.186 PSNR 22.4 | SSIM 0.76

Figure 11: Additional 100-shot objective-ablation cases for user 199 (left) and user 26 (right). Within each case, panels show (a) Original, (b) the Expert C reference, (c) the Generic Enhancement Prior, (d) the Full objective, (e) without Lpref , (f) without Lgap , (g) without the fidelity-plus-gap regularization group, and (h) without scorer guidance. Scores are frozen personalized scorer outputs, while PSNR and SSIM use Expert C as reference. On both displayed images, Full improves the frozen-scorer output and Expert C fidelity over the Generic Enhancement Prior, whereas removing the full regularization group causes severe overexposure. tions. FedPAIE does not require a personalized target image for any photograph in the local enhancement set. A new user provides a small private rated support set to calibrate the scorer, while the enhancer is adapted on a separate set of ordinary unpaired local photographs. Generic paired retouches are used only to learn the common CLUT prior and do not encode the target style of the new user. This decomposition replaces user-specific retouch supervision with scorer-mediated preference supervision while keeping the user’s raw photographs and ratings local.

I.4

Federated Preference and Visual Learning

Federated learning trains a shared model through client-side optimization and server aggregation without collecting raw client samples (McMahan et al. 2017). Extensions address statistical heterogeneity and non-IID client distributions (Li et al. 2020; Mohri, Sivek, and Suresh 2019). Federated collaborative filtering, matrix factorization, and recommendation learn user or item representations from decentralized interactions (Ammad-ud-din et al. 2019; Chai et al. 2021; Liang, Pan, and Ming 2021; Yi et al. 2021; Liu et al. 2023). Their outputs support ranking and recommendation rather than image formation. Federated vision research has demonstrated decentralized learning for object detection, general computer-vision tasks, personalized aesthetic assessment, and parameter-efficient video moderation (Liu et al. 2020; He et al. 2021; Xiong, Yu, and Shen 2023; Tao et al. 2025). The closest task to FedPAIE is federated personalized image-aesthetics assessment (Xiong, Yu, and Shen 2023), which predicts userspecific aesthetic scores from decentralized data. As an assessment method, its output remains a prediction rather than a user-specific enhanced image. The other federated vision systems demonstrate decentralized visual learning across deployed, benchmark, and parameter-efficient settings, but do not learn an individual’s color-grading preference or translate it into a personalized photo transformation. FedPAIE crosses this task boundary. Federated training supplies a shared aesthetic scorer initialization, local support

ratings calibrate it to a new user, and the frozen personalized scorer subsequently guides a local enhancer on unpaired images. Only scorer parameters and scalar aggregation counts participate in the federated exchange. The generic enhancer is learned independently from generic paired retouches, and user-specific enhancer adaptation remains on device. This division limits the federated component to preference learning instead of communicating a full image-to-image model. As stated in the main paper, this design preserves raw-data locality by keeping photos, ratings, and personalized models on device while communicating only lightweight scorer updates and scalar aggregation counts.

I.5

Positioning and Technical Novelty of FedPAIE

Tab. 27 summarizes the functional boundary between the closest personalized-aesthetics task settings. FedPAIE introduces an end-to-end federated personalization formulation that connects Federated Aesthetic Preference Learning, Personalized Scorer Calibration, and Frozen-Scorer-Guided Enhancer Adaptation. This formulation directly converts decentralized sparse ratings into user-specific 3D-LUT transformations on unpaired local images. From Decentralized Ratings to Image Transformations. Previous federated preference methods terminate at a score, ranking, or recommendation, while personalized enhancement methods directly learn a transformation from centrally available preference or retouch supervision. FedPAIE connects these two endpoints through a calibrated scorer that is both user-specific and differentiable. This connection enables sparse ratings to guide an enhancer even when the local photographs have no corresponding personalized targets. Separation of Shared Priors and Private Adaptation. FedPAIE learns two independent shared initializations. The aesthetic scorer is federated across decentralized ratings, while the Generic Enhancement Prior is initialized from a pretrained checkpoint obtained from ordinary paired retouches. Personalization then occurs entirely on device in a fixed order: calibrate the scorer from the private rated support set, freeze it, and adapt the enhancer from unpaired photographs.

Method

Output / inference

Preference signal

Federated PIAA (Xiong, Yu, and Shen 2023)

Personalized aesthetic score

Decentralized user image–rating data

PAA+ (Zhong et al. 2025)

Personalized aesthetic score

Surveys, scores, and accumulated multimodal feedback

PPS (Kim, Yoo, and Kim 2026)

Personalized Pairwise user judgments over photographic-style output style candidates

FedPAIE

Personalized enhanced image 0.293M enhancer at inference

Sparse scalar ratings and unpaired local photos

Fed. training

Raw-data locality

Pers. transform

Transform supervision / local adaptation

Yes

Yes (reported)

No

N/A

Not addressed

Not reported

No

N/A

Not addressed

Not reported

Yes

Pairwise style candidates Unpaired local adaptation not reported

Yes (scorer)

Yes (protocol)

Yes

No user-specific retouch target Unpaired local photos

Table 27: Method-level comparison of the closest personalized-aesthetics settings. “Not reported” means that the cited work does not state the corresponding raw-data-locality or local-adaptation property. “Not addressed” marks a capability outside its stated task, and N/A marks a transformation-specific field that does not apply to a score-only method. Generic and LUT-based enhancement methods remain complementary architecture context in Sec. I.1 rather than being mixed into this task-setting comparison. This separation keeps the user’s preference evidence local and avoids treating a population retouching style as the user’s target style. Support-Aware and Stability-Aware Personalization. The Support-Dependent Scorer Mask restricts the adaptable parameter blocks when only 10 ratings are available and permits broader calibration with 100 ratings. The local ranking and variance objectives complement regression by preserving relative preference information and score dispersion. During enhancer adaptation, the scorer remains fixed and the preference objective is balanced by pixel, perceptual, and excess-gap regularization. The resulting design addresses two distinct risks of few-shot scorer guidance: overfitting the preference model and over-optimizing its imperfect output. A Lightweight Deployment Boundary. Personalization updates only the CNN backbone and coefficient head of the enhancer while keeping the compressed LUT bases fixed. The scorer is a training-time component and is removed after adaptation. Inference therefore retains only the 0.293Mparameter personalized enhancer, which distinguishes FedPAIE from scorer-only federated aesthetics methods and from personalization pipelines that retain a larger preference or style model at deployment. Taken together, these properties establish FedPAIE as a distinct federated personalization framework. To the best of our knowledge, FedPAIE is the first federated method for personalized aesthetic image enhancement and color grading. It transforms decentralized sparse ratings into lightweight user-specific color grading on unpaired local images while keeping raw photos and ratings on device.

Input

Global

Client 194

Client 199

Client 200

Client 204

Client 210

Figure 12: Randomly selected shared-input comparison for 100-shot Clients 194, 199, 200, 204, and 210. Each row contains a FiveK input, the Generic Enhancement Prior output, and five personalized outputs. Python’s random.Random(42) selects one of three pre-generated, non-overlapping groups. Stable client-dependent color transformations remain visible across the shared inputs. Brightness Change

Contrast Change

Saturation Change

0.10 0.05

0.100

0.3

0.06

0.075

0.00

0.2

0.050

0.04

−0.05

0.1

−0.10

4 8 3 3 2 al 10 10 12 15 17 ob Gl nt nt nt nt nt ie ie ie ie ie Cl Cl Cl Cl Cl

Colorfulness Change 0.125

0.4

0.08

0.025

0.02

0.0

0.000

0.00

−0.1

−0.025

al

ob

Gl

4

nt

ie

Cl

8

10

nt

ie

Cl

10

nt

ie

Cl

3 3 2 12 15 17 nt nt ie ie Cl Cl

4 8 3 3 2 al 10 10 12 15 17 ob nt nt nt nt nt ie ie ie ie ie Cl Cl Cl Cl Cl

Gl

4 8 3 3 2 al 10 10 12 15 17 ob nt nt nt nt nt ie ie ie ie ie Cl Cl Cl Cl Cl

Gl

Figure 13: Attribute changes from the input for the Generic Enhancement Prior and five main-paper clients. Bars and error bars show the mean and variation across qualitative examples. The statistics characterize each client’s enhancement style.

Record · ID 414081 · SHA-256 6cd37f2f2ab018c9
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.