Conceptio › Archive › arXiv CS
arXiv CSopen access

Large-Scale High-Quality 3D Gaussian Head Reconstruction from Multi-View Captures

2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
neural-networks
machine learning, deep learning, neural networks

arXiv:2605.04035v1 [cs.CV] 5 May 2026

Large-Scale High-Quality 3D Gaussian Head Reconstruction from Multi-View Captures Evangelos Ntavelis, Sean Wu† , Mohamad Shahbazi, Fabio Maninchedda, Dmitry Kostiaev, Artem Sevastopolsky, Vittorio Megaro, Trevor Phillips, Alejandro Blumentals, Shridhar Ravikumar, Mehak Gupta, Reinhard Knothe, Jeronimo Bayer, Matthias Vestner, Simon Schaefer, Thomas Etterlin, Christian Zimmermann, Mathias Deschler, Peter Kaufmann, Stefan Brugger, Sebastian Martin, Brian Amberg, and Tom Runia Apple Abstract. We propose HeadsUp, a scalable feed-forward method for reconstructing high-quality 3D Gaussian heads from large-scale multicamera setups. Our method employs an efficient encoder-decoder architecture that compresses input views into a compact latent representation. This latent representation is then decoded into a set of UV-parameterized 3D Gaussians anchored to a neutral head template. This UV representation decouples the number of 3D Gaussians from the number and resolution of input images, enabling training with many high-resolution input views. We train and evaluate our model on an internal dataset with more than 10 000 subjects, which is an order of magnitude larger than existing multi-view human head datasets. HeadsUp achieves state-of-the-art reconstruction quality and generalizes to novel identities without test-time optimization. We extensively analyze the scaling behavior of our model across identities, views, and model capacity, revealing practical insights for quality-compute trade-offs. Finally, we highlight the strength of our latent space by showcasing two downstream applications: generating novel 3D identities and animating the 3D heads with expression blendshapes. Keywords: 3D Reconstruction · 3D Gaussian Splatting · Digital Humans

1

Introduction

High-fidelity 3D head assets are foundational to photorealistic digital humans, enabling convincing and authentic social co-presence in immersive digital environments. Such assets are particularly valuable for view-consistent, close-up rendering in applications such as telepresence, actor digitization, and content creation [40]. To achieve this level of quality, multi-camera capture systems have become a standard approach for producing dense calibrated images of human heads [12]. However, it remains a challenge to reliably and efficiently transform these detailed captures into compact 3D reconstructions at scale [38, 82]. Existing solutions for converting multi-view head captures into renderable assets span a spectrum that reflects a recurring tension between reconstruction †

Work done during internship at Apple.

2

E. Ntavelis et al.

Large-scale Pretraining

Finetuning on Ava-256

Renders from Validation Subjects

Renders from Validation Subjects

Fig. 1. We introduce HeadsUp, a novel feed-forward approach leveraging 3D Gaussians to predict high-quality avatars. By scaling to thousands of subjects and diverse expressions, our method achieves exceptional rendering quality on completely held-out subjects. Notice the accurate, high-resolution recovery of intricate fine details, such as eyelashes, complex earrings, teeth and tongue. The figure displays renders from novel subjects unseen during training.

fidelity and throughput. At one end, instance-specific optimization methods, such as per-subject fitting using Neural Radiance Fields [4, 37, 54] or 3D Gaussian Splatting (3DGS) [31,63], serve as strong baselines for high-quality view-consistent rendering. However, the computational cost of per-identity optimization makes their large-scale deployment challenging. At the other end, recent feed-forward reconstruction methods [6, 8, 38, 91] amortize computation across datasets and enable fast inference, but their compute and memory costs typically scale with the number and resolution of input views. This hinders the full utilization of dense, high-resolution capture setups. Orthogonally, another line of work targets animatable avatars by predicting geometry and appearance in a canonical, modelconditioned space [9, 32, 82]. While this supports temporal coherence and control, it often sacrifices the representation capacity needed for high-quality rendering. Together, these trends suggest a clear need for methods that can fully leverage multi-camera rigs (i.e., many views, high resolution, many identities) while producing compact head assets suitable for photorealistic rendering. To address this need, we prioritize high-fidelity reconstruction quality from high-resolution multi-view captures spanning thousands of identities and millions of frames. We propose HeadsUp, a scalable feed-forward approach for reconstructing high-fidelity 3D Gaussian head assets from large-scale multi-camera studio captures. Specifically, our method transforms input multi-view images to the UV parameterization of a set of 3D Gaussians attached to a neutral head template. In contrast to pixel-aligned prediction [27, 38], the UV-space formulation decouples the number of output Gaussians from the number and resolution of input images. This allows our method to scale gracefully with the number of high-resolution input views, which is crucial for resolving fine details such as hair strands and jewelry. Our feed-forward model is based on an encoder-decoder architecture: a cross-attention transformer efficiently encodes the input images into a compact 2D latent representation, which is then decoded into the corresponding Gaussian UV maps. To better capture the high-frequency details, we introduce a background

HeadsUp: Large-scale Gaussian Head Reconstruction

3

model to improve fine boundary structures such as hair, and a high-resolution finetuning stage that leverages multi-scale and region-specific losses to enhance overall render quality and the fidelity of core facial features like the eyes and mouth. To demonstrate the reconstruction quality and scalability of our method, we train and evaluate our model on a large-scale internal dataset that is an order of magnitude larger than existing multi-view face datasets, comprising over 10 000 subjects with diverse appearances and facial expressions. For comparison with prior work, we also report results on the publicly available Ava-256 dataset [53]. HeadsUp achieves state-of-the-art reconstruction quality on novel identities and expressions (examples in Fig. 1). We further characterize the scaling behavior of our model with respect to the number of training subjects and input views, as well as the model’s capacity. Finally, we highlight the strength of our learned latent space by showcasing two downstream applications: generating novel 3D identities with latent diffusion models and animating the 3D heads from expression blendshapes. In summary, our main contributions are as follows: – We address the challenge of large-scale, high-fidelity 3D head reconstruction from multi-camera studio captures spanning thousands of identities and millions of frames. – We introduce HeadsUp, a feed-forward model that predicts UV maps that encode the parameters of 3D Gaussians anchored to a neutral head template. Our model decouples the output representation from the input resolution and view count, enabling efficient scaling on dense, high-resolution data. – We introduce explicit background modeling and a tailored training strategy to improve reconstruction in high-frequency regions (hair, eyes, and mouth). – We achieve state-of-the-art reconstruction quality on novel identities and expressions, and characterize the scaling behavior of our model across identities, views, and model capacity.

2

Related Work

Novel View Synthesis of Human Heads. Various 3D representations have been explored for head reconstruction and novel view synthesis. Earlier attempts based on Neural Radiance Fields [2, 17, 37, 57, 95] have succeeded at photorealistic rendering but required lengthy per-scene optimization. Implicit surface methods [93] and point-based approaches [90, 94] offered alternatives with different trade-offs between quality and efficiency. Hybrid methods like MonoNPHM [18] combine neural parametric models with implicit representations for monocular reconstruction. Works on Codec Avatars [42, 48, 52, 63] and others [71] achieve high-fidelity head avatars through learning on multiple dense multi-view captures. Often, these methods involve per-subject optimization or fine-tuning at test time. 3DGS optimization-based approaches have quickly become popular due to their efficiency and more explicit representation, albeit limited in scalability

4

E. Ntavelis et al.

because of the large number of Gaussians required. GaussianAvatars [59] ties 3D Gaussians to a parametric model [45] for fully controllable heads but requires minutes of optimization per identity, excluding multi-view tracking time. SplattingAvatar [64] embeds Gaussians within triangle meshes for hybrid mesh-Gaussian avatars. GAF [69] distills multi-view diffusion priors using pseudo-ground truths to enhance monocular avatar reconstruction. These works, along with [70,77] and others, demonstrate that 3DGS is an excellent representation for detailed head avatars. However, the reliance on slow optimization makes scaling to thousands of subjects with high-resolution studio captures challenging. Feed-forward 3D Head Reconstruction. Following the advances in Large Reconstruction Models (LRMs) [6, 8, 24, 68, 75, 78, 87], recent works develop feed-forward 3DGS head models to bypass per-subject optimization. Multi-view feed-forward methods address the well-constrained setting of reconstruction from multiple input images. GPAvatar [10] learns efficient Gaussian projections from multi-view inputs by employing deep feature extractors. HeadGAP [92] learns generalizable Gaussian priors from multi-view data and personalizes to new identities from few images, though it still requires a persubject adaptation step. Avat3r [38] regresses animatable 3D head avatars from as few as four images using DUSt3R position maps [76] and Sapiens features [33]. While template-free, its pixel-aligned Gaussians scale linearly with input views, causing memory constraints and potential temporal inconsistencies. The reliance on DUSt3R and Sapiens make the pipeline time- and memory-constrained in a practical setting. Similarly, the concurrent method FastGHA [27] achieves few-shot real-time animation via pixel-aligned features, yet faces comparable scaling limitations. Finally, Pippo [30] achieves high quality results by employing a diffusion transformer trained on a large-scale dataset. While practical for single images, inference scales to minutes for multi-view captures, and strict 3D view consistency is not guaranteed. Parallel works on full-body reconstruction explore similar concepts, such as efficiently reconstructing animatable humans from posefree images [60, 61], regressing Gaussians in UV space [39], or training DiT-based generators [81]. Despite achieving impressive full-body results, these methods remain bottlenecked by the limited facial resolution of available datasets. Single-view feed-forward methods tackle the ill-posed problem of reconstructing human heads from a single input image. GAGAvatar [9] is one of the first generalizable triplane-based Gaussian head models. Portrait4D [14, 15] learns novel view synthesis and driving from synthetic 2D data and pseudo multiview supervision, generated from a triplane-based generator. However, triplane representations often struggle with extreme novel viewpoints and are sensitive to image cropping. LAM [22] and PanoLAM [43] perform one-shot head synthesis based on UV-aligned Gaussians, similar to our method but with a focus on rigging. FastAvatar [46] predicts residual Gaussians from a canonical template in <10 ms, specifically targeting the fast rigging setting. PercHead [56] uses perceptual supervision from DINOv2 [55] and SAM [35] for disentangled geometry and appearance control. Our method partially borrows inspirations from some of these methods and uses a UV-aligned Gaussian map that naturally supports

HeadsUp: Large-scale Gaussian Head Reconstruction

5

arbitrary number of views, including single view, while remaining fast and supporting a wide range of novel viewpoints. Rather than optimizing for realtime rigging, in this work we prioritize maximum reconstruction quality. Like Pippo [30], several recent methods rely on diffusion priors to reconstruct human heads [7, 44, 50, 51, 79, 80, 84–86], but this reliance comes at computational cost and compromises rendering fidelity. 3D GAN inversion represents a parallel track for single-image 3D head synthesis. PanoHead [1] and SphereHead [41] extend EG3D [5] to 360◦ head synthesis with back-of-head supervision. Encoder-based latent inversion methods like TriPlaneNet [3] and GOAE [83], alongside optimization-based PTI [62], facilitate reconstruction from real images, while DiffPortrait3D [20] and VOODOO3D [72] focus on one-shot reenactment. InvertAvatar [89] adapts this inversion strategy to multiple images for incremental improvement. However, triplanes can fundamentally limit the level of achievable resolution and expressivity. Furthermore, triplane-based generators rely on external 2D head pose estimation that can introduce compounding errors; these methods also usually impose strict cropping requirements on the input image and camera assumptions [5, 66]. Summary. Ultimately, our work distinguishes itself from relevant prior work by combining several key properties: (1) a fully feed-forward architecture that avoids lengthy optimization to a novel subject and decouples memory from the input view count, enabling massive scalability; (2) universal support for monocular, sparse, or dense multi-view inputs; (3) independence from strict parametric face models, allowing for natural reconstruction of accessories and diverse expressions; and (4) state-of-the-art photorealism and subject recognizability in novel views.

3

Method

Here we introduce our method HeadsUp. Fig. 2 shows the overall architecture of our model. Given N calibrated time-synchronized images, our model jointly predicts 3D Gaussians for the foreground and background. Explicit background modeling bypasses the need for foreground matting, which often fails to accurately segment high-frequency boundary details like hair or jewelry. Our model is based on UV-parameterized 3D Gaussians anchored to a template mesh. Our feed-forward network consists of two main components: a multi-view encoder based on a transformer architecture to extract latents from the input images, and a 3D Gaussian decoder to map the latents to the 3D Gaussians’ UV parameters. To ensure high-fidelity rendering of core facial features without prohibitive computational costs, we employ a two-stage training strategy where a high-resolution finetuning stage leverages region-specific and multi-scale supervision. Below, we detail our model architecture and training methodology. 3.1

UV-Parameterized 3D Gaussian Splatting

Our model predicts a multi-channel 2D UV map, representing a fixed number of Gaussians anchored to a template mesh. The channels of the UV map correspond

6

E. Ntavelis et al. N Input Views

Foreground Model Novel Views

Image Tokens Ffg NxHxWx512

FG ResNet Encoder

K

V

K

V

K

V

Learnable Tokens Q

...

Transformer Block

Q

Transformer Block

Q

Transformer Block

Latents Z

1x64x64x512

Foreground Positional Prior

Background Model

FG ResNet Decoder

Rendered Image

1x1x1x32

Ground Truth

BG Gaussians

Background Latent zbg BG ResNet Encoder

FG Gaussians

1x64x64x512

BG ResNet Decoder

Fig. 2. Overview of HeadsUp. Our method reconstructs high-fidelity 3D Gaussian heads from multi-view images. Given a set of input views, our model utilizes a transformer-based encoder and a 3D Gaussian decoder to predict UV-parameterized 3D Gaussians for both the foreground and background. The model is trained end-to-end using a combination of photometric and perceptual supervision.

to the Gaussian attributes: position µ ∈ R3 , scale s ∈ R3+ , quaternion rotation q ∈ R4 , opacity α ∈ [0, 1], and spherical harmonic color coefficients {c(ℓ,m) } up to degree L = 1. The foreground Gaussians are anchored to a neutral head template shared among all identities in a canonical coordinate system. The head canonical coordinate system is defined with its origin at the mid-pupil point of the template and its orientation aligned to the Frankfurt plane [26]. The background Gaussians are anchored to a sphere template fitted to the capture rig and are transformed into the canonical coordinate system for joint rendering. The UV formulation in our method offers several advantages, including: (1) Shared geometric prior. The canonical mesh topology provides consistent spatial structure across subjects and expressions, allowing the network to focus on appearance and local variations rather than coarse global head geometry. (2) Efficient multi-view aggregation. The mesh-anchored UV parameterization decouples the number of output Gaussians from the number and resolution of input images, allowing for efficient aggregation of information from dense and high-resolution captures. (3) Robustness to tracking errors. As the 3D Gaussians are anchored to a fixed neutral template, our representation only requires rigid head pose tracking and does not rely on fragile and error-prone facial expression tracking for inference. 𛲜

3.2

Network Architecture

Here we describe our network architecture, composed of a multi-view encoder and a Gaussian UV decoder.

HeadsUp: Large-scale Gaussian Head Reconstruction

7

3×H×W Multi-View Encoder. Given N calibrated input views {Ii }N , i=1 , Ii ∈ R N 4×4 and their camera extrinsics {Ei }i=1 , Ei ∈ R , along with intrinsics {Ki }N , i=1 Ki ∈ R3×3 , we first patchify the images and convert each into patch embeddings eip ∈ Rd×h×w . To explicitly encode the camera geometry into the features, we concatenate per-patch Plücker embeddings [65] with the corresponding patch embeddings along the channel dimension: \mathbf {f}^{i}_p = \text {Concat}(\mathbf {e}^{i}_p, \mathcal {P}(\mathbf {K}_i, \mathbf {E}_i)) \quad \in \mathbb {R}^{(d+6) \times h \times w}

(1)

where P(·) computes the 6D Plücker coordinates. Two parallel convolutional encoders process these features to disentangle each input view into foreground c×hf ×wf c′ ×hf ×wf and background features, Ffg and Fbg , respectively. i ∈R i ∈R We employ a transformer architecture [16, 74] to map these unstructured multi-view foreground features to a low-resolution 2D latent representation of the target 3D head. The transformer converts a 2D grid of learnable tokens Q ∈ Rdz ×hz ×wz to the 2D latent Z ∈ Rdz ×hz ×wz by aggregating information from multi-view features through cross-attention layers: \mathbf {Z} = \text {CrossAttnTransformer}(\mathbf {Q}, \mathbf {F}_{\text {fg}}).

(2)

When discussing our latent representation throughout the paper, we refer to the foreground latent Z. To model the background, we use a shallow convolutional network with global N average pooling to map the stacked multi-view background features {Fbg i }i=1 to dbg the background latent zbg ∈ R . This design choice is motivated by the fact that the background is mostly constant between different frames, apart from small variations such as lighting. 3D Gaussian UV Decoder. We use a convolutional decoder to convert the resulting latent variable Z into a high-resolution UV feature map U ∈ RH×W ×23 . Specifically, for each UV location (u, v), the output UV features are mapped to the corresponding 3D Gaussian attributes as follows: \boldsymbol {\mu }_{u,v} &= \mathbf {V}{(u,v)} + \mathbf {U}^{(\mu )}(u,v), \quad \|\mathbf {U}^{(\mu )}(u,v)\| \leq \delta _{\max } \\ \mathbf {s}_{u,v} &= \exp (\mathbf {U}^{(s)}(u,v))\\ \mathbf {q}_{u,v} &= \text {normalize}(\mathbf {U}^{(q)}(u,v))\\ \alpha _{u,v} &= \sigma (\mathbf {U}^{(\alpha )}(u,v)) \\ \{\mathbf {c}_{u,v}^{(\ell ,m)}\} &= \mathbf {U}^{(c)}(u,v),

(7) where V(u, v) is the corresponding 3D vertex position on the template mesh V ∈ R(H×W )×3 , and δmax is the position offset bound (empirically set to 200mm). Similarly, zbg is decoded into the background 3D Gaussians using a similar but separate decoder (with δmax = 10mm). During training, we find it important to perform a warm-up period of 1000 iterations, where opacity and scale attributes are detached from the gradient backpropagation graph.

8

E. Ntavelis et al.

3.3

Training Objectives

Our overall training loss consists of reconstruction and regularization terms: \mathcal {L}_{\mathrm {total}} = \lambda _{\mathrm {L1}} \mathcal {L}_{\mathrm {L1}} + \lambda _{\mathrm {LPIPS}} \mathcal {L}_{\mathrm {LPIPS}} + \lambda _{\mathrm {adv}} \mathcal {L}_{\mathrm {adv}} + \lambda _{\mathrm {pos}} \mathcal {L}_{\mathrm {pos}} + \lambda _{\mathrm {mask}} \mathcal {L}_{\mathrm {mask}} + \lambda _{\mathrm {TV}} \mathcal {L}_{\mathrm {TV}}, (8) where the λ hyperparameters control the relative influence of each term. Reconstruction Losses. Let I and Igt denote the rendered and ground-truth composite images (foreground and background) from a randomly sampled view. We optimize an L1 photometric loss LL1 = ∥I − Igt ∥1 and a multi-scale LPIPS perceptual loss LLPIPS [29]. To enhance high-frequency details, we employ a Perceptual Discriminator [19, 67] to compute an adversarial loss Ladv . To maintain training stability, Ladv is activated only after 240k iterations, and the discriminator operates on random 256 × 256 spatial crops of the input. Regularization Terms. To guide the geometry during the warm-up phase, we utilize expression-tracked meshes to regularize the Gaussian positions. Specifically, the loss term Lpos penalizes the distance between µ(u, v) and the 3D point Ve (u, v), derived via barycentric interpolation of the Gaussian UV coordinates on a tracked mesh Ve . Concurrently, a silhouette loss Lmask minimizes the discrepancy between the rendered foreground alpha map and the ground-truth segmentation mask Mgt . The weights for both Lpos and Lmask are decayed over the course of training. Finally, a Total Variation loss LTV is applied to the rendered UV-space colors [36] to encourage spatial smoothness and prevent surface holes. 3.4

Two-Stage Training

Training our model on a large dataset of high-resolution input images can be computationally expensive as the transformer scales quadratically with the number of input tokens. Therefore, we adopt a two-stage training strategy: we first train on 2× downscaled images, followed by a high-resolution finetuning stage that uses the native image resolution. During high-resolution finetuning we make two key modifications: – Region-Specific Losses. We introduce region-specific losses for areas with high-frequency details. Specifically, we extract image crops around the eyes and mouth using the head canonical coordinate frame and supervise with additional multi-scale LPIPS perceptual losses. – Multi-Resolution Loss Strategy. We observed that applying global perceptual and discriminator losses directly at native resolution leads to training instability, likely due to noisy high-frequency gradients. Therefore, to ensure stable convergence, we compute these losses on 2× downsampled outputs. More implementation details are provided in the Supplementary Material.

HeadsUp: Large-scale Gaussian Head Reconstruction

4

Experiments

4.1

Datasets

9

Internal Multi-View Head Dataset. Our internal dataset contains over 10 000 unique participants recorded in a head-focused rig with 16 calibrated RGB cameras. The images are of the resolution 1000 × 750. Participants were asked to perform various facial expressions and speech sequences. We train on 10 000 subjects with 100 frames per subject sampled for expression diversity. We evaluate on 100 frames from 50 validation subjects. All participants provided written informed consent for use of their data. The internal dataset will not be publicly released. Images shown in this paper are from subjects who provided explicit written consent for the use of their images in the publication and visualizations. Ava-256 Dataset. We also finetune and evaluate our model on the public dataset Ava-256 [53], which contains high-resolution multi-view head captures. We use the 4TB version of the dataset, containing 256 subjects recorded with 80 RGB cameras. The images are of the resolution 1024 × 667. Following Avat3r’s setup [38], we train on 244 subjects with 1000 frames per subject and evaluate on 12 validation subjects with around 2000 sampled frames in total. Unlike previous methods restricted to frontal views, we sample views from all 80 cameras.

4.2

Experimental Setup

Training. We train our model on the internal dataset for 900K steps with 2× downscaled images (batch size 64) in the first stage, followed by 200K steps of full-resolution finetuning (batch size 32). We use 10 input images for this setup. Training with the Adam optimizer [34] with a learning rate of 2 × 10−4 converges in approximately 10 days on 16 H100 GPUs with bfloat16 precision. We perform full-resolution finetuning on Ava-256 for 200K steps with 16 input images, converging in less than a day. Metrics. We evaluate rendering quality with three image metrics: Peak Signalto-Noise Ratio (PSNR), Structural Similarity (SSIM), and Learned Perceptual Image Patch Similarity (LPIPS) [88]. Following Avat3r [38], we also report two face-specific metrics: Average Keypoint Distance (AKD) which measures the distance in pixels between 2D keypoints estimated from PIPNet [28], and cosine similarity (CSIM) of ArcFace identity embeddings [13].

4.3

Baseline

For our baseline comparison, we evaluate our method against Avat3r, representing the state-of-the-art in feed-forward 3D head reconstruction.

10

E. Ntavelis et al.

N =4 Input Avat3r Ours

Input

N =6 Avat3r Ours

N = 16 Input Ours

GT

Fig. 3. Visual Comparison on Ava-256. Our method produces sharper reconstructions with better identity preservation compared to prior work. Increasing the number of views permits reconstruction of details like earrings, hair and skin texture. Additionally, our background model successfully captures intricate head-boundary details that previous foreground-masking techniques typically discard; we only use the background model during training. Due to GPU memory constraints during training, Avat3r is limited to a maximum of 6 views.

Avat3r [38]. Avat3r uses DUSt3R [76] position maps and Sapiens features [33] to reconstruct head avatars from sparse views. Since the official code is not released, we carefully reimplemented Avat3r and made some modifications to enable fair comparison to our work. Specifically, we adapt Avat3r from a sparseview self-reenactment setting to a large-scale 3D head reconstruction pipeline by (1) removing the expression rigging module and (2) providing time-synchronized multi-view images as input. Furthermore, the original method precomputes DUSt3R and Sapiens features for only 10 frames per subject, whereas our experimental setup scales this to 1000 frames per subject. As extracting DUSt3R position maps at this scale is computationally prohibitive, we instead compute them using the more efficient VGGT [75]. 4.4

Baseline Comparison

We compare our method with Avat3r [38] on our internal dataset (Internal10K) and Ava-256 dataset [53] using different metrics. In Ava-256 experiments, both our model and Avat3r are finetuned from the models pretrained on Internal10K. As shown in Tab. 1, our method significantly outperforms the baseline, with major gains in rendering quality (PSNR, LPIPS) and face fidelity (AKD, CSIM) metrics, while requiring more than an order of magnitude fewer Gaussians. Qualitative results in Fig. 3 also show renders with much higher fidelity and sharper details

HeadsUp: Large-scale Gaussian Head Reconstruction

Internal10K #V #G PSNR↑ SSIM↑ LPIPS↓ AKD↓ CSIM↑ Avat3r

4 0.8M

24.10

0.830

0.371

12.78

0.738

Avat3r

6 1.1M

24.37

0.831

0.362

12.99

0.787

Ours

10 65K

29.25

0.821

0.117

10.39 0.922

Ava-256 #V #G PSNR↑ SSIM↑ LPIPS↓ AKD↓ CSIM↑ Avat3r

4 0.8M

21.54

0.788

0.370

6.39

Avat3r

6 1.1M

22.31

0.795

0.359

6.10

0.738 0.775

Ours

4 65K

23.64

0.805

0.178

5.09

0.833

Ours

6 65K

24.62

0.815

0.161

4.70

0.873

Ours

16 65K

26.13

0.831

0.110

4.27

0.914

11

Table 1. Quantitative Comparison. We evaluate both our method and Avat3r [38] on two datasets: our Internal10K dataset, and Ava-256 [53] (finetuned from the Internal10K models), across varying numbers of input views (#V). Our approach significantly outperforms the baseline with major gains in rendering quality (PSNR, LPIPS) and facial fidelity (AKD, CSIM) metrics. Notably, we achieve these improvements while requiring substantially fewer 3D Gaussians (#G). Also, Avat3r is bottlenecked by GPU memory limits at 6 views on our training GPU budget.

for our method, capturing fine hair and facial details, with sharper eye and mouth regions than the baselines. Notably, Avat3r is bottle-necked by GPU memory limits at 6 views, whereas our efficient formulation allows for many more inputs. More results and video visualizations are provided in the Supplementary Material. 4.5

Analysis and Ablation Study

We ablate the proposed components of our overall method in Fig. 4. We conduct extensive analysis to validate our design choices and characterize the scaling behavior of our method. We organize our analysis around five key aspects: number of training identities, number of input views, representation capacity, number of target views used for supervision, and our high-resolution finetuning strategy. Moreover, we study the effect of the background model, region-specific perceptual losses, and multi-resolution loss. Number of training identities. As depicted in Fig. 4a, PSNR improves 1.7–1.8 dB per doubling of subjects up to 2K, with reduced performance gains after 4K. Identity preservation (CSIM) scales more steeply, from 0.441 (250 subjects) to 0.888 (8K). Models trained on <1K subjects fail catastrophically on out-of-distribution faces. Visual comparisons are provided in Fig. 5. Number of input views. We train our model with an increasing number of random input views. For evaluation, we use a fixed set of views with varying size per experiment, e.g. the frontal view for the monocular case. As shown in Fig. 4b and Fig. 6, more views improve the quality, with diminishing gains after 8 views. Notably, our monocular model is able to convincingly reconstruct the subject. Latent and Gaussian UV resolution. As shown in Fig. 4c, increasing the latent resolution yields more significant improvements than the UV resolution.

12

E. Ntavelis et al. (a) Training Subjects

(b) Input Views

(c) Latent vs. UV Res.

(g) Background model ablation

30 0.10

No Background Model

With Background Model 1

28

26

2

1

2

1

0.20 24 PSNR LPIPS

250

500

1K

2K

# Training Subjects

(d) Mesh Type 30

LPIPS PSNR

0.10

10K 1

0.094

27

0.16

26.77

29.54

29.28

29.2 29.0

29.000.095

Expr.-Tracked

Fixed Neutral

32

64

32

3

LPIPS PSNR w/o Eye Loss w/o Mouth Loss w/o HalfRes Perceptual Full Model

0.05

1

3

3

3

30

0.09

28

0.15

26

0.20

24

0.25

0.10

0.096

2

128

Latent Resolution

0.10

4

4

4

4

0.10

0.18 28.6

0.184

16 16

0.09

28.89

28.8 26

8

2

0.25

(f) HighRes Finetuning

0.093 0.09

0.09

0.12 29.6 0.14

4

# Input Views

LPIPS PSNR

29.8

29.4 28

2

(e) Target Views

30.0

0.096

28.89

29

4K

PSNR UV 256 PSNR UV 512 LPIPS UV 256 LPIPS UV 512

PSNR LPIPS

LPIPS

22

PSNR

1

LPIPS

PSNR

0.15

0.10 2

4

# Target Views

8

Full Image

Eye Crop

Mouth Crop

Fig. 4. Ablation study on a single-stage model trained for 500K steps with 10K subjects, 10 input views, 32×32 latent and 256×256 Gaussian UV resolution unless stated otherwise. (a) Training data scaling: log-linear improvement up to 2K subjects, diminishing returns beyond 4K. (b) Input view scaling: quality improves with more views, with diminishing returns after 8. (c) Model capacity: increasing latent resolution yields larger gains than increasing Gaussian UV resolution. (d) Template mesh type: a fixed neutral mesh outperforms expression-tracked meshes. (e) Number of target views: more supervision views improve geometric consistency (f ) Highresolution finetuning: Effect of different components in our second-stage training strategy. (g) Background model ablation: Our background modeling permits reconstruction of foreground details such as strands of hair, without background artifacts caused by imperfect image matting techniques.

Template mesh type. Our ablation in Fig. 4d shows that using a fixed neutral template significantly outperforms using an expression-tracked one. Number of target views. Adding more supervision views improves all the metrics by encouraging better view-consistency as shown in Fig. 4e. High-resolution finetuning. As discussed in Sec. 3.4, after our low-resolution training is converged, we finetune our model on high-resolution target and input images. In Fig. 4f we examine our design choices for the second training stage: Region-specific losses. Our eye and mouth crop losses improve reconstruction quality in critical facial features where human perception is most sensitive. Finetuning perceptual losses. Downsampling the rendered and real inputs to Discriminator and LPIPS losses is important for the stability of the adversarial losses. Naively computing LPIPS on high resolution images leads to quality regressions. Background Gaussian model. As shown in Fig. 4g, removing the background model degrades rendering quality for regions with semi-transparent areas or fine foreground elements. Without explicit background modeling, the model hallucinates background artifacts, such as discoloration in the hair. Since computed foreground masks are not pixel-perfect and view-consistent, background details

HeadsUp: Large-scale Gaussian Head Reconstruction 250

500

1K

4K

10K

13

GT

Fig. 5. Training data scaling. Models trained on fewer subjects fail to generalize to reconstruction of novel identities. At 250 subjects, facial features and hair color deviate significantly from ground truth. The reconstruction quality improves with more training data. On this validation set, the quality improves marginally after 4K subjects.

N =1

N =2

N =4

N =6

N =8

N = 16

GT

Fig. 6. Impact of the number of input views. Reconstruction quality scales naturally with the number of input images. A single frontal view (N = 1) yields blurry results, identity drift, and fails to recover shirt text. However, adding more views progressively resolves these ambiguities, yielding clear improvements in fine details like the teeth and hair.

are often combined with foreground elements within masked ground truth images used for supervision. By eliminating the need for these masks, our dedicated background model prevents artifacts from background regions, such as the multicamera capture rig, and allows the foreground model to focus more Gaussians on facial details.

E. Ntavelis et al.

"A young boy with a

"A photo of an old

eyes, blonde hair"

gold chain"

blue shirt, green

man, dark skin,

"A woman with

"A young woman a

"An Asian woman

hair, yellow dress"

earrings"

and a white dress"

brown eyes, purple

nose piercing and

with red lipstick

Source identity

14

Source expression

(a) Text-driven identity generation

(b) Blendshape-driven latent animation

Fig. 7. Downstream Applications. (a) Text-driven identity generation: Novel identities generated by a text-conditioned diffusion model trained on our latents. Using nearest-neighbor face similarity, we verify that these synthesized subjects do not exist in our training set. (b) Blendshape-driven latent animation: We train a network conditioned on expression blendshapes, applying a target expression (blue box) to a source identity (green box). The model successfully animates the faces while preserving the subject’s appearance. Both applications operate entirely within the latent space, requiring no per-subject fine-tuning.

4.6

Downstream Applications

Text-driven Identity Generation. Our compact, yet information-rich latent space enables a range of downstream applications, such as novel identity generation. To show this, we train a text-conditioned DiT [58] on a large dataset of latents Z precomputed from our base model. At inference time, we sample latents and decode them into Gaussians using our frozen decoder. Fig. 7a shows randomly sampled identities from our trained model. Based on face-similarity analysis, we have confirmed these identities do not appear in the training set. Blendshape-driven Latent Animation. We showcase facial animation controlled by expression blendshapes, operating entirely within our latent space. Given triplets (Zn , Zb , b) of a neutral latent Zn , expression latent Zb of the same subject and the corresponding blendshape coefficients b, we train a transformer Fθ to predict the target expression latent Ẑb = Fθ (Zn , b) with supervised losses on the latents, Gaussians and renders. Results are displayed in Fig. 7b.

5

Conclusion

We present HeadsUp, a highly scalable feed-forward approach for state-of-the-art reconstruction of 3D Gaussian heads from multi-camera studio captures. By anchoring a compact set of 3D Gaussians to a neutral UV template and employing a lightweight cross-attention transformer, our method achieves photorealistic rendering while gracefully scaling with respect to the number of input views. Crucially, our explicit background modeling eliminates the reliance on imperfect

HeadsUp: Large-scale Gaussian Head Reconstruction

15

segmentation masks. Combined with a two-stage training strategy and regionspecific perceptual losses, this enables the faithful reconstruction of challenging, high-frequency details such as hair strands and jewelry. We show that HeadsUp successfully scales to a production-level dataset of 10 000 unique identities, delivering state-of-the-art reconstruction quality and exhibiting robust generalization to unseen subjects. Finally, beyond feed-forward reconstruction, we show that our learned latent space enables downstream generative applications, including the text-driven synthesis of novel identities and blendshape-driven latent animation. Additional details on the methodology and experimental setup, along with qualitative results including rendered images and videos, are provided in the supplementary material.

Acknowledgements We thank Simon Biland, Armin Kappeler, Alexey Artemov and Rick Zhang for their support and valuable feedback.

References 1. An, S., Xu, H., Shi, Y., Song, G., Ogras, U.Y., Luo, L.: PanoHead: Geometry-Aware 3D Full-Head Synthesis in 360°. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 20950–20959 (2023) 2. Athar, S., Xu, Z., Sunkavalli, K., Shechtman, E., Shu, Z.: RigNeRF: Fully Controllable Neural 3D Portraits. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 20364–20373 (2022) 3. Bhattarai, A.R., Nießner, M., Sevastopolsky, A.: TriPlaneNet: An Encoder for EG3D Inversion. In: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV). pp. 3533–3542 (2024) 4. Bühler, M.C., Li, G., Wood, E., Helminger, L., Chen, X., Shah, T., Wang, D., Garbin, S., Orts-Escolano, S., Hilliges, O., Lagun, D., Riviere, J., Gotardo, P., Beeler, T., Meka, A., Sarkar, K.: Cafca: High-quality Novel View Synthesis of Expressive Faces from Casual Few-shot Captures. In: SIGGRAPH Asia 2024 Conference Papers. ACM (2024). https://doi.org/10.1145/3680528.3687580 5. Chan, E.R., Lin, C.Z., Chan, M.A., Nagano, K., Pan, B., De Mello, S., Gallo, O., Guibas, L.J., Tremblay, J., Khamis, S., et al.: Efficient Geometry-aware 3D Generative Adversarial Networks. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 16123–16133 (2022) 6. Charatan, D., Li, S., Tagliasacchi, A., Sitzmann, V.: pixelSplat: 3D Gaussian Splats from Image Pairs for Scalable Generalizable 3D Reconstruction. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 19457–19467 (2024) 7. Chen, J., Li, C., Zhang, J., Zhu, L., Huang, B., Chen, H., Lee, G.H.: Generalizable Human Gaussians from Single-View Image. arXiv preprint arXiv:2406.06050 (2024) 8. Chen, Y., Xu, H., Zheng, C., Zhuang, B., Pollefeys, M., Geiger, A., Cham, T.J., Cai, J.: MVSplat: Efficient 3D Gaussian Splatting from Sparse Multi-View Images. In: European Conference on Computer Vision (ECCV). pp. 370–386. Springer (2024)

16

E. Ntavelis et al.

9. Chu, X., Harada, T.: Generalizable and Animatable Gaussian Head Avatar. NeurIPS (2024) 10. Chu, X., Li, Y., Zeng, A., Yang, T., Lin, L., Liu, Y., Harada, T.: GPAvatar: Generalizable and Precise Head Avatar from Image(s). arXiv preprint arXiv:2401.10215 (2024) 11. Chung, H.W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, Y., Wang, X., Dehghani, M., Brahma, S., et al.: Scaling instruction-finetuned language models. Journal of Machine Learning Research 25(70), 1–53 (2024) 12. Debevec, P., Hawkins, T., Tchou, C., Duiker, H.P., Sarokin, W.: Acquiring the Reflectance Field of a Human Face. In: SIGGRAPH. New Orleans, LA (Jul 2000), http://ict.usc.edu/pubs/Acquiring%20the%20Re%EF%AC%82ectance%20Field% 20of%20a%20Human%20Face.pdf 13. Deng, J., Guo, J., Xue, N., Zafeiriou, S.: ArcFace: Additive Angular Margin Loss for Deep Face Recognition. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 4690–4699 (2019) 14. Deng, Y., Wang, D., Ren, X., Chen, X., Wang, B.: Portrait4D: Learning One-Shot 4D Head Avatar Synthesis using Synthetic Data. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2024) 15. Deng, Y., Wang, D., Wang, B.: Portrait4D-v2: Pseudo Multi-View Data Creates Better 4D Head Synthesizer. In: European Conference on Computer Vision (ECCV). pp. 303–321. Springer (2024) 16. Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., Houlsby, N.: An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. ICLR (2021) 17. Gafni, G., Thies, J., Zollhöfer, M., Nießner, M.: Dynamic Neural Radiance Fields for Monocular 4D Facial Avatar Reconstruction. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 8649–8658 (2021) 18. Giebenhain, S., Kirschstein, T., Georgopoulos, M., Rünz, M., Agapito, L., Nießner, M.: MonoNPHM: Dynamic Head Reconstruction from Monocular Videos. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 21051–21061 (2024) 19. Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., Bengio, Y.: Generative Adversarial Networks. In: Advances in neural information processing systems. pp. 2672–2680 (2014), http://papers.nips.cc/ paper/5423-generative-adversarial-nets.pdf 20. Gu, Y., Xu, H., Xie, Y., Song, G., Shi, Y., Di, Y., Ye, P., et al.: DiffPortrait3D: Controllable Diffusion for Zero-Shot Portrait View Synthesis. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2024) 21. He, K., Zhang, X., Ren, S., Sun, J.: Deep Residual Learning for Image Recognition. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 770–778 (2016) 22. He, Y., Gu, X., Ye, X., Xu, C., Zhao, Z., Dong, Y., Yuan, W., Dong, Z., Bo, L.: LAM: Large Avatar Model for One-shot Animatable Gaussian Head. In: Proceedings of the Special Interest Group on Computer Graphics and Interactive Techniques Conference Conference Papers. pp. 1–13 (2025) 23. Ho, J., Salimans, T.: Classifier-free diffusion guidance. In: NeurIPS 2021 Workshop on Deep Generative Models and Downstream Applications (2021), https: //openreview.net/forum?id=qw8AKxfYbI

HeadsUp: Large-scale Gaussian Head Reconstruction

17

24. Hong, Y., Zhang, K., Gu, J., Bi, S., Zhou, Y., Liu, D., Liu, F., Sunkavalli, K., Bui, T., Tan, H.: LRM: Large Reconstruction Model for Single Image to 3D. arXiv preprint arXiv:2311.04400 (2023) 25. Hoogeboom, E., Mensink, T., Heek, J., Lamerigts, K., Gao, R., Salimans, T.: Simpler diffusion: 1.5 fid on imagenet512 with pixel-space diffusion. In: Proceedings of the Computer Vision and Pattern Recognition Conference. pp. 18062–18071 (2025) 26. International Organization for Standardization: ISO 7250-1:2017 Basic human body measurements for technological design — Part 1: Body measurement definitions and landmarks (2017), https://www.iso.org/standard/65246.html, accessed: 2026-02-16 27. Ji, X., Weiss, S., Kansy, M., Naruniec, J., Cao, X., Solenthaler, B., Bradley, D.: FastGHA: Generalized Few-Shot 3D Gaussian Head Avatars with Real-Time Animation. arXiv preprint arXiv:2601.13837 (2026) 28. Jin, H., Liao, S., Shao, L.: Pixel-in-Pixel Net: Towards Efficient Facial Landmark Detection in the Wild. International Journal of Computer Vision 129(12), 3174–3194 (2021) 29. Johnson, J., Alahi, A., Fei-Fei, L.: Perceptual Losses for Real-Time Style Transfer and Super-Resolution. In: European Conference on Computer Vision (2016) 30. Kant, Y., Weber, E., Kim, J.K., Khirodkar, R., Zhaoen, S., Martinez, J., Gilitschenski, I., Saito, S., Bagautdinov, T.: Pippo: High-Resolution Multi-View Humans from a Single Image. In: CVPR. pp. 16418–16429 (2025) 31. Kerbl, B., Kopanas, G., Leimkühler, T., Drettakis, G.: 3D Gaussian Splatting for Real-Time Radiance Field Rendering. ACM TOG 42(4) (July 2023), https: //repo-sam.inria.fr/fungraph/3d-gaussian-splatting/ 32. Khakhulin, T., Sklyarova, V., Lempitsky, V., Zakharov, E.: Realistic One-shot Mesh-based Head Avatars. arXiv preprint arXiv:2206.08343 (2022) 33. Khirodkar, R., Bagautdinov, T., Martinez, J., Zhaoen, S., James, A., Selednik, P., Anderson, S., Saito, S.: Sapiens: Foundation for Human Vision Models. In: ECCV. pp. 206–228. Springer (2024) 34. Kingma, D.P.: Adam: A Method for Stochastic Optimization. arXiv preprint arXiv:1412.6980 (2014) 35. Kirillov, A., Mintun, E., Ravi, N., et al.: Segment Anything. arXiv preprint arXiv:2304.02643 (2023) 36. Kirschstein, T., Giebenhain, S., Tang, J., Georgopoulos, M., Nießner, M.: GGHead: Fast and Generalizable 3D Gaussian Heads. In: SIGGRAPH Asia 2024 Conference Papers. pp. 1–11 (2024) 37. Kirschstein, T., Qian, S., Giebenhain, S., Walter, T., Nießner, M.: NeRSemble: Multi-view Radiance Field Reconstruction of Human Heads. ACM TOG 42(4), 1–14 (2023) 38. Kirschstein, T., Romero, J., Sevastopolsky, A., Nießner, M., Saito, S.: Avat3r: Large Animatable Gaussian Reconstruction Model for High-fidelity 3D Head Avatars. arXiv preprint arXiv:2502.20220 (2025) 39. Kwon, Y., Fang, B., Lu, Y., Dong, H., Zhang, C., Vicente Carrasco, F., MosellaMontoro, A., Xu, J., Takagi, S., Kim, D., Prakash, A., De la Torre, F.: Generalizable Human Gaussians for Sparse View Synthesis. In: Proceedings of the European Conference on Computer Vision (ECCV) (2024) 40. Lawrence, J., Goldman, D.B., Achar, S., Blascovich, G.M., Desloge, J.G., Fortes, T., Gomez, E.M., Häberling, S., Hoppe, H., Huibers, A., Knaus, C., Kuschak, B., Martin-Brualla, R., Nover, H., Russell, A.I., Seitz, S.M., Tong, K.: Project Starline:

18

E. Ntavelis et al.

A High-Fidelity Telepresence System. ACM Transactions on Graphics (Proc. of SIGGRAPH Asia) 40(6) (2021) 41. Li, H., Chen, C., Shi, T., Qiu, Y., An, S., Chen, G., et al.: SphereHead: Stable 3D Full-head Synthesis with Spherical Tri-plane Representation. In: European Conference on Computer Vision (ECCV). Springer (2024) 42. Li, J., Cao, C., Schwartz, G., Khirodkar, R., Saito, S., et al.: URAvatar: Universal Relightable Gaussian Codec Avatars. In: SIGGRAPH Asia 2024 Conference Papers. ACM (2024) 43. Li, P., He, Y., Hu, Y., Dong, Y., Yuan, W., Liu, Y., Zhu, S., Cheng, G., Dong, Z., Guo, Y.: PanoLAM: Large Avatar Model for Gaussian Full-Head Synthesis from One-shot Unposed Image. arXiv preprint arXiv:2509.07552 (2025) 44. Li, P., Zheng, W., Liu, Y., Yu, T., Li, Y., Qi, X., Chi, X., Xia, S., Cao, Y.P., Xue, W., et al.: PSHuman: Photorealistic Single-Image 3D Human Reconstruction Using Cross-Scale Multiview Diffusion and Explicit Remeshing. In: Proceedings of the computer vision and pattern recognition conference. pp. 16008–16018 (2025) 45. Li, T., Bolkart, T., Black, M.J., Li, H., Romero, J.: Learning a Model of Facial Shape and Expression from 4D Scans. ACM TOG 36(6), 194–1 (2017) 46. Liang, H., Ge, Z., Tiwari, A., Majee, S., Godaliyadda, G.M.D., Veeraraghavan, A., Balakrishnan, G.: FastAvatar: Instant 3D Gaussian Splatting for Faces from Single Unconstrained Poses. arXiv preprint arXiv:2508.18389 (2025) 47. Lin, S., Ryabtsev, A., Sengupta, S., Curless, B.L., Seitz, S.M., KemelmacherShlizerman, I.: Real-Time High-Resolution Background Matting. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 8762–8771 (2021) 48. Lombardi, S., Saragih, J., Simon, T., Sheikh, Y.: Deep Appearance Models for Face Rendering. ACM Transactions on Graphics (TOG) 37(4), 1–13 (2018) 49. Lu, C., Zhou, Y., Bao, F., Chen, J., LI, C., Zhu, J.: Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps. In: Koyejo, S., Mohamed, S., Agarwal, A., Belgrave, D., Cho, K., Oh, A. (eds.) Advances in Neural Information Processing Systems. vol. 35, pp. 5775–5787. Curran Associates, Inc. (2022), https://proceedings.neurips.cc/paper_files/paper/2022/file/ 260a14acce2a89dad36adc8eefe7c59e-Paper-Conference.pdf 50. Lu, Y., Dong, J., Kwon, Y., Zhao, Q., Dai, B., De la Torre, F.: GAS: Generative Avatar Synthesis from a Single Image. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 12883–12893 (2025) 51. Lyu, W., Zhou, Y., Yang, M.H., Shu, Z.: FaceLift: Learning Generalizable Single Image 3D Face Reconstruction from Synthetic Heads. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 12691–12701 (2025) 52. Ma, S., Simon, T., Saragih, J., Wang, D., Li, Y., De La Torre, F., Sheikh, Y.: Pixel Codec Avatars. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 64–73 (2021) 53. Martinez, J., Kim, E., Romero, J., Bagautdinov, T., Saito, S., Yu, S.I., Anderson, S., Zollhöfer, M., Wang, T.L., Bai, S., Li, C., Wei, S.E., Joshi, R., Borsos, W., Simon, T., Saragih, J., Theodosis, P., Greene, A., Josyula, A., Maeta, S.M., Jewett, A.I., Venshtain, S., Heilman, C., Chen, Y.T., Fu, S., Elshaer, M.E.A., Du, T., Wu, L., Chen, S.C., Kang, K., Wu, M., Emad, Y., Longay, S., Brewer, A., Shah, H., Booth, J., Koska, T., Haidle, K., Andromalos, M., Hsu, J., Dauer, T., Selednik, P., Godisart, T., Ardisson, S., Cipperly, M., Humberston, B., Farr, L., Hansen, B., Guo, P., Braun, D., Krenn, S., Wen, H., Evans, L., Fadeeva, N., Stewart, M., Schwartz, G., Gupta, D., Moon, G., Guo, K., Dong, Y., Xu, Y., Shiratori, T., Prada, F.,

HeadsUp: Large-scale Gaussian Head Reconstruction

19

Pires, B.R., Peng, B., Buffalini, J., Trimble, A., McPhail, K., Schoeller, M., Sheikh, Y.: Codec Avatar Studio: Paired Human Captures for Complete, Driveable, and Generalizable Avatars. NeurIPS (2024) 54. Mildenhall, B., Srinivasan, P.P., Tancik, M., Barron, J.T., Ramamoorthi, R., Ng, R.: NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis. In: European Conference on Computer Vision (ECCV) (2020) 55. Oquab, M., Darcet, T., Moutakanni, T., Vo, H., Szafraniec, M., Khalidov, V., Fernandez, P., Haziza, D., Massa, F., El-Nouby, A., et al.: DINOv2: Learning Robust Visual Features without Supervision. arXiv preprint arXiv:2304.07193 (2023) 56. Oroz, A., Nießner, M., Kirschstein, T.: PercHead: Perceptual Head Model for Single-Image 3D Head Reconstruction & Editing. arXiv preprint arXiv:2511.02777 (2025) 57. Park, K., Sinha, U., Barron, J.T., Bouaziz, S., Goldman, D.B., Seitz, S.M., MartinBrualla, R.: Nerfies: Deformable Neural Radiance Fields. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). pp. 5865–5874 (2021) 58. Peebles, W., Xie, S.: Scalable Diffusion Models with Transformers. In: Proceedings of the IEEE/CVF international conference on computer vision. pp. 4195–4205 (2023) 59. Qian, S., Kirschstein, T., Schoneveld, L., Davoli, D., Giebenhain, S., Nießner, M.: GaussianAvatars: Photorealistic Head Avatars with Rigged 3D Gaussians. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2024) 60. Qiu, L., Gu, X., Li, P., Zuo, Q., Shen, W., Zhang, J., Qiu, K., Yuan, W., Chen, G., Dong, Z., Bo, L.: LHM: Large Animatable Human Reconstruction Model from a Single Image in Seconds. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) (2025) 61. Qiu, L., Li, P., Zuo, Q., Gu, X., Dong, Y., Yuan, W., Zhu, S., Han, X., Chen, G., Dong, Z.: PF-LHM: 3D Animatable Avatar Reconstruction from Pose-free Articulated Human Images. arXiv preprint arXiv:2506.13766 (2025) 62. Roich, D., Mokady, R., Bermano, A.H., Cohen-Or, D.: Pivotal Tuning for Latentbased Editing of Real Images. ACM Transactions on Graphics (TOG) 42(1), 1–13 (2022) 63. Saito, S., Schwartz, G., Simon, T., Li, J., Nam, G.: Relightable Gaussian Codec Avatars. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 130–141 (2024) 64. Shao, Z., Wang, Z., Li, Z., Wang, D., Lin, X., Zhang, Y., Fan, M., Wang, Z.: SplattingAvatar: Realistic Real-Time Human Avatars with Mesh-Embedded Gaussian Splatting. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 16922–16932 (2024) 65. Sitzmann, V., Rezchikov, S., Freeman, W.T., Tenenbaum, J.B., Durand, F.: Light Field Networks: Neural Scene Representations with Single-Evaluation Rendering. In: Proc. NeurIPS (2021) 66. Skorokhodov, I., Siarohin, A., Xu, Y., Ren, J., Lee, H.Y., Wonka, P., Tulyakov, S.: 3D Generation on ImageNet. arXiv preprint arXiv:2303.01416 (2023) 67. Sungatullina, D., Zakharov, E., Ulyanov, D., Lempitsky, V.: Image Manipulation with Perceptual Discriminators. In: Proceedings of the European Conference on Computer Vision (ECCV) (September 2018)

20

E. Ntavelis et al.

68. Szymanowicz, S., Rupprecht, C., Vedaldi, A.: Splatter Image: Ultra-Fast Single-View 3D Reconstruction. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 10208–10217 (2024) 69. Tang, J., Davoli, D., Kirschstein, T., Schoneveld, L., Niessner, M.: GAF: Gaussian Avatar Reconstruction from Monocular Videos via Multi-view Diffusion. In: Proceedings of the Computer Vision and Pattern Recognition Conference. pp. 5546–5558 (2025) 70. Teotia, K., Kim, H., Garrido, P., Habermann, M., Elgharib, M., Theobalt, C.: GaussianHeads: End-to-End Learning of Drivable Gaussian Head Avatars from Coarse-to-fine Representations. ACM Transactions on Graphics (SIGGRAPH Asia) 43(6) (2024) 71. Teotia, K., Rhodin, H., Mendiratta, M., Kim, H., Habermann, M., Theobalt, C.: Audio Driven Universal Gaussian Head Avatars. In: Proceedings of the SIGGRAPH Asia 2025 Conference Papers. pp. 1–12 (2025) 72. Tran, P., Zakharov, E., Ho, L.N., Tran, A.T., Hu, L., Li, H.: VOODOO 3D: Volumetric Portrait Disentanglement for One-Shot 3D Head Reenactment. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2024) 73. Umeyama, S.: Least-squares estimation of transformation parameters between two point patterns. IEEE Transactions on Pattern Analysis and Machine Intelligence 13(4), 376–380 (1991). https://doi.org/10.1109/34.88573 74. Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, L., Polosukhin, I.: Attention Is All You Need. In: Guyon, I., Luxburg, U.V., Bengio, S., Wallach, H., Fergus, R., Vishwanathan, S., Garnett, R. (eds.) Advances in Neural Information Processing Systems. vol. 30. Curran Associates, Inc. (2017), https://proceedings.neurips.cc/paper_files/paper/2017/file/ 3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf 75. Wang, J., Chen, M., Karaev, N., Vedaldi, A., Rupprecht, C., Novotny, D.: VGGT: Visual Geometry Grounded Transformer. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2025) 76. Wang, S., Leroy, V., Cabon, Y., Chidlovskii, B., Revaud, J.: DUSt3R: Geometric 3D Vision Made Easy. In: CVPR. pp. 20697–20709 (2024) 77. Xiang, J., Gao, X., Guo, Y., Zhang, J.: FlashAvatar: High-fidelity head avatar with efficient gaussian embedding. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 1802–1812 (2024) 78. Xu, Y., Shi, Z., Yifan, W., Chen, H., Yang, C., Peng, S., Shen, Y., Wetzstein, G.: GRM: Large Gaussian Reconstruction Model for Efficient 3D Reconstruction and Generation. In: European Conference on Computer Vision. pp. 1–20. Springer (2024) 79. Xue, Y., Xie, X., Marin, R., Pons-Moll, G.: Human-3Diffusion: Realistic Avatar Creation via Explicit 3D Consistent Diffusion Models. Advances in Neural Information Processing Systems 37, 99601–99645 (2024) 80. Yang, J., Wu, T., Fogarty, K., Zhong, F., Oztireli, C.: PSHead: 3D Head Reconstruction from a Single Image with Diffusion Prior and Self-Enhancement. In: Computer Graphics Forum. p. e70279. Wiley Online Library (2025) 81. Yang, Y., Liu, F., Lu, Y., Zhao, Q., Wu, P., Zhai, W., Yi, R., Cao, Y., Ma, L., Zha, Z.J., et al.: SIGMAN: Scaling 3D Human Gaussian Generation with Millions of Assets. ICCV (2025) 82. Ye, Z., Zhong, T., Ren, Y., Yang, J., Li, W., Huang, J., Jiang, Z., He, J., Huang, R., Liu, J., Zhang, C., Yin, X., Ma, Z., Zhao, Z.: Real3D-Portrait: One-shot

HeadsUp: Large-scale Gaussian Head Reconstruction

21

Realistic 3D Talking Portrait Synthesis. In: International Conference on Learning Representations (ICLR) (2024), https://openreview.net/forum?id=7ERQPyR2eb, spotlight 83. Yuan, Z., Zhu, Y., Li, Y., Liu, H., Yuan, C.: Make Encoder Great Again in 3D GAN Inversion through Geometry and Occlusion-Aware Encoding. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). pp. 2437–2447 (2023) 84. Zhang, B., Cheng, Y., Wang, C., Zhang, T., Yang, J., Tang, Y., Zhao, F., Chen, D., Guo, B.: RodinHD: High-Fidelity 3D Avatar Generation with Diffusion Models. In: European Conference on Computer Vision. pp. 465–483. Springer (2024) 85. Zhang, J., Gao, Y., Zhan, J., Wang, W., Zhang, Y., Zhao, H., Zhang, L.: HighQuality 3D Head Reconstruction from Any Single Portrait Image. arXiv preprint arXiv:2503.08516 (2025) 86. Zhang, J., Li, X., Zhang, Q., Cao, Y., Shan, Y., Liao, J.: HumanRef: Single Image to 3D Human Generation via Reference-Guided Diffusion. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 1844–1854 (2024) 87. Zhang, K., Bi, S., Tan, H., Xiangli, Y., Zhao, N., Sunkavalli, K., Xu, Z.: GS-LRM: Large Reconstruction Model for 3D Gaussian Splatting. In: Proceedings of the European Conference on Computer Vision (ECCV) (2024) 88. Zhang, R., Isola, P., Efros, A.A., Shechtman, E., Wang, O.: The Unreasonable Effectiveness of Deep Features as a Perceptual Metric. In: CVPR. pp. 586–595 (2018) 89. Zhao, X., Sun, J., Wang, L., Suo, J., Liu, Y.: InvertAvatar: Incremental GAN Inversion for Generalized Head Avatars. In: ACM SIGGRAPH 2024 Conference Papers. pp. 1–10 (2024) 90. Zhao, Z., Bao, Z., Li, Q., Qiu, G., Liu, K.: PSAvatar: A Point-based Shape Model for Real-Time Head Avatar Animation with 3D Gaussian Splatting. arXiv preprint arXiv:2401.12900 (2024) 91. Zheng, S., Tu, H., Zhou, B., Shao, R., Liu, B., Zhang, S., Nie, L., Liu, Y.: GPSGaussian: Generalizable Pixel-wise 3D Gaussian Splatting for Real-time Human Novel View Synthesis. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2024) 92. Zheng, X., Wen, C., Li, Z., Zhang, W., Su, Z., Chang, X., Zhao, Y., Lv, Z., Zhang, X., Zhang, Y., et al.: HeadGAP: Few-shot 3D head avatar via generalizable gaussian priors. In: 2025 International Conference on 3D Vision (3DV). pp. 946–957. IEEE (2025) 93. Zheng, Y., Abrevaya, V.F., Bühler, M.C., Chen, X., Black, M.J., Hilliges, O.: I M Avatar: Implicit Morphable Head Avatars from Videos. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 13545–13555 (2022) 94. Zheng, Y., Yifan, W., Wetzstein, G., Black, M.J., Hilliges, O.: PointAvatar: Deformable Point-based Head Avatars from Videos. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 21057–21067 (2023) 95. Zielonka, W., Bolkart, T., Thies, J.: Instant volumetric head avatars. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 4574– 4584 (2023)

22

E. Ntavelis et al.

Supplementary Material In this Supplementary Material, we provide additional results, implementation details, and information on our training procedures. We encourage the reader to view the HeadsUp videos on our webpage.

A

Inference Speed

We analyze the reconstruction time and scalability of our approach compared to the Avat3r baseline given a varying number of input images, as detailed in Tab. A.1. Because the baseline architecture struggles with the dense aggregation of multiview features, their processing time increases drastically as more views are added. Avat3r operates at sub-second frame rates for 4 and 6 views, and completely fails due to Out-of-Memory (OOM) errors during training when the input exceeds 6 views. Note that in the reported results for Avat3r, while the Sapiens [33] feature maps are computed on-the-fly, the VGGT [75] position maps are computed offline. Conversely, our method processes spatial features much more efficiently. While our reconstruction time naturally scales with the number of input views, it remains orders of magnitude faster than the baseline. Furthermore, our efficient memory management ensures that the model easily accommodates 10 or more views during both training and inference without exhausting GPU memory, allowing for higher-fidelity reconstructions without the computational bottleneck.

B

Additional Ablations

Table B.2 provides detailed numerical results for the ablation studies summarized in Fig. 4 of the main paper. All experiments use a single-stage model trained for 500K steps with 10K subjects, 10 input views, 32×32 latent and 256×256 Gaussian UV resolution unless stated otherwise. Number of Training Subjects. We study the effect of training data scale in Tab. B.2a. All metrics improve log-linearly as the number of subjects increases from 250 to 2K, with PSNR rising by over 5 dB across this range. Beyond 4K subjects, gains begin to saturate on this validation dataset: increasing from 4K to 10K yields a 0.43 dB improvement in PSNR. Table A.1. Inference speed comparison on a single A100 GPU for predicting 3D Gaussians from multi-view images. Our method is more than an order of magnitude faster than Avat3r, enabling the 3D reconstruction of large-scale datasets. Note that in the reported results for Avat3r, the VGGT position maps are computed offline.

Avat3r Ours

4 Views

FPS ↑ 6 Views

16 Views

0.18 7.14

0.09 4.94

OOM 2.93

Reconstruction Time (s) ↓ 4 Views 6 Views 16 Views 5.6 0.14

10.8 0.20

OOM 0.33

HeadsUp: Large-scale Gaussian Head Reconstruction

N =2

N =4

N =6

N =8

N = 16

Out of

Out of

Memory

Memory

GT

Out of

Out of

Memory

Memory

Ours

Avat3r

Ours

Avat3r

N =1

23

Fig. B.1. Ours vs. Avat3r: Impact of the number of input views. Avat3r runs out of memory for N > 6 input views, while our method scales to 16 or more views. When comparing on the same number of input views, our method outperforms Avat3r with sharper reconstructions and better identity preservation. All methods were trained for 500K steps with 10K subjects at 500 × 375 resolution.

Number of Input Views. Fig. B.1 compares how our method and Avat3r scale with the number of input views. Avat3r’s reliance on heavy foundation models (DUSt3R [76], Sapiens [33]) and its per-pixel Gaussian prediction incur substantial memory overhead that grows quadratically with the number of image tokens. This limits Avat3r to at most N =6 views on a single A100/H100 GPU (80 GB VRAM, batch size 1) before exceeding memory during training. Avat3r’s reconstructions are blurry and lack high-frequency details because (1) it relies on predicted point maps that introduce geometric discontinuities (as noted in FastGHA [27]), and (2) its confidence-based masking can aggressively remove valid foreground regions. In contrast, our architecture decouples the output Gaussian count from the input resolution and view count through the UV parameterization (discussed in Sec. 3.1 of the main paper), enabling efficient scaling to N =16 views and beyond with minimal per-view memory overhead. As shown in Fig. B.1, our method produces consistently sharper reconstructions with better-preserved identity details at every view count, including the monocular setting. We use N =10 views

24

E. Ntavelis et al.

for the Internal10k dataset, as quality saturates beyond this point (see Tab. B.2b in the Supplementary Material and Fig. 4b in the main paper). For Ava256, we use N =16 views, as we find it provides a good trade-off of viewpoint coverage and training time. Mesh Type. We compare using a fixed neutral mesh versus an expressiontracked mesh as the UV parameterization substrate in Tab. B.2c. The fixed neutral mesh outperforms expression-tracked meshes by a large margin across all metrics (+2.12 dB PSNR, −0.088 LPIPS, −0.81 AKD). We attribute this to the fact that expression-tracked meshes introduce noisy per-frame vertex displacements that the model must account for, whereas a fixed neutral mesh provides a stable canonical surface that allows the Gaussian decoder to focus entirely on modeling appearance variation. Vertex Loss. We ablate the vertex position regularization loss (Lpos ) in Tab. B.2d. Although incorporating Lpos results in meaningful quantitative improvements, particularly in PSNR and LPIPS, this regularization primarily serves to stabilize early training and accelerate warm-up. The model converges to a similar visual quality without explicit vertex supervision across all metrics, demonstrating that our use of tracked meshes is a lightweight prior for enhanced training stability rather than a fundamental limitation of the proposed method. Latent and Gaussian UV Resolution. Tab. B.2e disentangles the effect of latent resolution (controlling model capacity) from Gaussian UV resolution (controlling the number of output Gaussians). Increasing the latent size from 16×16 to 128×128 yields consistent improvements across all metrics, with PSNR rising from 27.59 to 29.66 dB. In contrast, doubling the Gaussian UV resolution from 256 to 512 at any fixed latent size provides only marginal gains. This indicates that model capacity, rather than Gaussian count, is the primary bottleneck for reconstruction quality in our architecture. Number of Target Views. We vary the number of target views used for supervision during training in Tab. B.2f. Increasing from 1 to 8 target views improves all metrics, with PSNR rising from 28.89 to 29.54 dB and AKD decreasing from 3.20 to 2.97. Supervising with more target views per training step enforces multi-view consistency and reduces geometric ambiguities, as the model must produce Gaussians that render correctly from diverse viewpoints simultaneously. Number of Transformer Blocks. Tab. B.2g ablates the number of transformer blocks in the encoder. Performance improves steadily from 2 to 8 blocks, with PSNR increasing from 28.24 to 28.89 dB. Adding further blocks beyond 8 yields no additional improvement: the 12-block model achieves identical PSNR and LPIPS while slightly degrading AKD. We therefore use 8 blocks as the default, balancing reconstruction quality with computational cost. High-Res Finetuning: Region-Specific Losses. Tab. B.2h evaluates the contribution of each component in our high-resolution finetuning stage. Removing the eye-region loss degrades eye crop PSNR by 1.56 dB, while removing the

HeadsUp: Large-scale Gaussian Head Reconstruction

25

mouth-region loss reduces mouth crop PSNR by 1.97 dB, confirming that regionspecific supervision is critical for faithfully reconstructing fine details in these perceptually important areas. Removing the half-resolution loss causes the largest overall degradation (−3.55 dB full-image PSNR), as it provides the coarse-to-fine gradient signal that stabilizes the high-resolution training. The full model achieves the best trade-off, with near-optimal performance across all regions.

C

Additional Visualizations

The supplementary webpage contains interactive video results organized into the following sections: – Ava-256 Validation Renders. Per-subject reconstruction results with multiexpression renderings from novel viewpoints, comparing our method against Avat3r across different encoder view counts. – Ava-256 Comparisons. Side-by-side rendering comparisons on held-out validation subjects, highlighting fine facial details such as pores, hair strands, and specular highlights. – Internal10K Renders. Results on our internal multi-camera capture dataset. – Blendshape Rigging. Interactive demonstrations of blendshape-driven deformation (neutral, smile, raised brows) showing consistent Gaussian field deformation across subjects. Additionally, in Fig. C.2, we provide more examples of generated novel identities discussed in Sec. 4.6 of the main paper.

D

Avat3r Baseline Reimplementation

As no official code is available, we reimplement Avat3r [38], adapting it from its original sparse-view self-reenactment setting (256 subjects, 10 input frames each) to our large-scale multi-view setup (10K+ subjects, 100 input frames each). Modifications. We make three key changes: (1) Removed expression rigging: We drop the cross-attention blocks for expression latents, as our setting provides time-synchronized multi-view images of the target expression. (2) Replaced DUSt3R with VGGT: Precomputing DUSt3R [76] position maps at our scale (10K subjects × 100 frames × 16 views) is prohibitively expensive. We use VGGT [75], which is faster and produces improved geometry. (3) Replaced Sapiens-2B with 1B: Precomputing Sapiens [33] requires 100 TB of storage for Internal10K alone. Instead, we compute Sapiens-1B on-the-fly during training. It performs comparably to 2B at half the compute cost. Position Map Geometry Prior. We align the normalized VGGT point maps to metric scale via a 7-DoF similarity transform [73]. This is computed per-frame using 2D facial landmarks projected onto the VGGT point cloud and their

26

E. Ntavelis et al.

Table B.2. Ablation studies on a single-stage model trained for 500K steps with 10K subjects, 10 input views, 32 × 32 latent and 256 × 256 Gaussian UV resolution unless mentioned otherwise. (a) Training data scaling: our model scales gracefully with the number of training subjects. (b) Input view scaling: quality improves with more input views, diminishing returns after 8. (c) Mesh type: a fixed neutral mesh outperforms expression-tracked meshes. (d) Vertex loss (evaluated at 360K steps): regularizing mesh vertices improves geometric consistency. (e) Model capacity: increasing the latent size yields better performance than increasing the number of Gaussians. (f ) Supervision density: more target views improve geometric consistency. (g) Transformer blocks: performance saturates at 8 blocks. (h) High-resolution finetuning: region-specific losses for eyes and mouth are critical. (a) Number of Training Subjects

250 500 1K 2K 4K 8K 10K

PSNR↑ LPIPS↓ AKD↓ 22.19 0.267 6.27 24.07 0.219 4.77 25.92 0.164 3.97 27.74 0.118 3.29 28.46 0.103 3.23 3.16 28.76 0.098 28.89 0.096 3.20

(b) Number of Input Views

1 2 4 8 16

PSNR↑ LPIPS↓ AKD↓ 22.98 0.208 5.08 24.37 0.174 4.75 26.51 0.135 3.91 28.26 0.105 3.35 29.03 0.096 3.36 (c) Mesh Type

PSNR↑ LPIPS↓ AKD↓ Expr.-Tracked 26.77 0.184 4.01 Fixed Neutral 28.89 0.096 3.20 (d) Vertex Loss

w/o Vertex λ w/ Vertex λ

PSNR↑ LPIPS↓ AKD↓ 27.68 0.114 3.65 28.46 0.100 3.31

(e) Latent / Gaussian UV Resolution

16/256 16/512 32/256 32/512 64/256 64/512 128/256 128/512

PSNR↑ LPIPS↓ AKD↓ 27.59 0.121 3.59 27.43 0.119 3.51 28.89 0.096 3.20 28.90 0.093 3.28 29.21 0.089 3.00 29.46 0.086 3.21 29.52 0.087 3.07 29.66 0.083 3.04

(f ) Number of Target Views

1 2 4 8

PSNR↑ LPIPS↓ AKD↓ 28.89 0.096 3.20 29.00 0.095 3.19 29.28 0.094 3.14 29.54 0.093 2.97

(g) Number of Transformer Blocks

2 4 6 8 12

PSNR↑ LPIPS↓ AKD↓ 28.24 0.102 3.40 28.51 0.102 3.22 28.80 0.099 3.27 28.89 0.096 3.20 28.89 0.096 3.36

(h) High-Res Finetuning: Region-Specific Losses

w/o Eye w/o Mouth w/o HalfRes Full Model

Full Image Eye Crop Mouth Crop PSNR↑ LPIPS↓ PSNR↑ LPIPS↓ PSNR↑ LPIPS↓ 28.33 0.191 30.38 0.093 32.27 0.081 28.20 0.188 31.78 0.069 30.17 0.118 24.80 0.267 28.26 0.122 30.54 0.118 28.35 0.190 31.94 0.070 32.14 0.083

HeadsUp: Large-scale Gaussian Head Reconstruction "A photo of a young man, blue eyes, red cheeks, necklace, white t-shirt"

"A photo of an elderly woman, silver hair, wrinkles, pearl earrings, knitted cardigan"

"A photo of a young woman, dark skin, short afro, gold earrings, yellow blouse"

"A photo of a man, "A photo of a woman, "A photo of a young "A photo of a man, "A photo of a man, Asian features, man, ginger hair, tan skin, green eyes, South olive skin, wavy heavy brow, crooked thick mustache, dark brown light freckles, scar on cheek, hair, dimples, square salt-and-pepper curly hair, blue jaw, green nose, leather jacket" red tank top" hair, denim vest" dress shirt" crew neck sweater"

27

"A photo of a man, broad shoulders, buzzcut, cleft chin, navy polo shirt"

Fig. C.2. Additional Generated Novel Identities. A diverse set of novel identities generated by a text-conditioned diffusion model trained on our learned head latents. Each row shows different generated subjects, demonstrating the diversity in age, gender, ethnicity, hairstyle, and facial features. These generated subjects do not exist in our training set, as confirmed by nearest-neighbor face similarity. Our supplementary webpage also provides rendered videos of these novel subjects.

corresponding 3D tracked mesh vertices. The aligned maps are transformed into the head canonical frame using the head pose. During the forward pass, we drop VGGT confidence map pixels below the 10th percentile to create a binary mask, which is applied to the predicted Gaussian positions as in Avat3r.

Encoder View Sampling. We adapt the two-step input view sampling of Avat3r by fixing Ncandidate = 16 diverse (not strictly frontal) cameras per frame during preprocessing to compute VGGT maps. During training, we uniformly subsample N =4 or N =6 views. This preserves viewpoint diversity while avoiding the cost of computing position maps for all possible view subsets. Unlike Avat3r we do not restrict the Ncandidate cameras to frontal-only cameras.

Training. We faithfully match Avat3r’s architecture, losses, and hyperparameters, adjusting only input resolution (full-resolution vs. 512 × 512 crops), batch size, and learning rate. Training proceeds in three stages: (1) 2× downsampled Internal10K (batch size 64, λlpips =0), (2) full-resolution Internal10K (batch size 16, λlpips =0.01), and (3) full-resolution Ava-256 (batch size 16, λlpips =0.01). Models are trained until validation metrics plateau on 16 H100 GPUs.

28

E. Ntavelis et al.

E

Implementation Details

E.1

Architectural Details

We provide detailed specifications for each component of our architecture. All hyperparameters are summarized in Tab. H.3. Foreground Encoder. Each input image Ii ∈ R3×H×W is patchified with a 7×7 convolution (stride 7) into patch embeddings of dimension d=256 (the image is first resized to be compatible with the patch size). Patch embeddings are then concatenated with 6-dimensional Plücker ray embeddings [65] encoding the camera geometry. A convolutional network with one downsampling stage (2 residual blocks at 512 channels) followed by 4 bottleneck residual blocks at 512 channels 512×hf ×wf maps the image patch tokens to foreground feature maps Ffg . i ∈R Background Encoder. A separate, lighter convolutional encoder produces 256×hf ×wf per-view background features Fbg , using two downsampling stages i ∈R (2 residual blocks at 128 channels and 2 residual blocks at 64 channels). These features are aggregated across views via global average pooling followed by a twolayer MLP, yielding a compact background latent zbg ∈ Rdbg . The background latent requires less capacity because the background is largely static across frames. Cross-Attention Transformer. The foreground features from all N input views are flattened into a set of key-value tokens. A transformer with 8 blocks, 8 attention heads, hidden dimension dz =512, and MLP dimension 1024 maps a 2D grid of hz ×wz =64×64 learnable query tokens to the foreground latent Z ∈ R512×64×64 via cross-attention. Each block applies layer normalization, multi-head cross-attention, and a feed-forward network with GELU activations. The query grid provides a fixed spatial structure that the decoder can directly reshape into a UV map, while cross-attention aggregates information from an arbitrary number of views without quadratic view-count scaling. Foreground Decoder. The latent Z is decoded into the Gaussian UV map U ∈ R256×256×23 via a pre-activation residual network [21]. The decoder applies two 2× nearest-neighbor upsampling stages with channel dimensions [512, 256], each containing two pre-activation residual blocks using 3×3 convolutional kernels and learned residual branch scaling. Nearest-neighbor upsampling avoids the checkerboard artifacts of transposed convolutions. A final 3×3 convolution projects to 32-dimensional per-texel features, from which the 23 Gaussian attributes are regressed. Following GRM [78], we apply sigmoid activations for opacities and L2 normalization for rotation quaternions. Position offsets use tanh scaled by δmax to bound displacements from the template mesh, scales use an exponential activation, and SH color coefficients are output directly without activation. This yields 256×256 ≈ 65K foreground Gaussians.

HeadsUp: Large-scale Gaussian Head Reconstruction

29

Background Decoder. The background latent zbg is decoded into a UV map of 512×512 ≈ 262K background Gaussians anchored to a sphere template fitted to the capture rig. Unlike the foreground decoder, the background decoder uses a residual network architecture with LeakyReLU activations and BatchNorm, progressively upsampling from 4×4 resolution through 7 stages with channel dimensions [512, 256, 128, 64, 32, 32, 32, 32]. A tighter position offset bound (δmax =10 mm vs. 200 mm for foreground) constrains Gaussians near the rig geometry. E.2

Multi-Scale Perceptual Loss

Our perceptual loss LLPIPS is implemented as a multi-scale LPIPS loss [88] using the official pretrained AlexNet-based LPIPS network. Rather than computing the perceptual similarity at a single resolution, we evaluate it at three scales to capture both fine detail and global structure: \mathcal {L}_{\mathrm {LPIPS}} = \sum _{k=0}^{2} \text {LPIPS}\!\left (\text {down}_{2^k}(I),\;\text {down}_{2^k}(I_{\mathrm {gt}})\right ),

(E.1)

where down2k denotes 2k × spatial downsampling via bilinear interpolation, so the three scales correspond to the native resolution (1×), 2× downsampled, and 4× downsampled. E.3

Downstream Applications

Here, we provide more details on the downstream applications we discussed in Sec. 4.6 of the main paper: Text-driven Identity Generation. To sample novel identities, we train a latent diffusion model that operates directly in the HeadsUp latent space. The HeadsUp encoder maps each identity to a latent tensor, which we treat as the diffusion target. Our denoising network is a DiT [58] with 10 transformer blocks, an embedding dimension of 512, an MLP hidden dimension of 2048. For textconditioned generation, we encode the input prompt using a frozen Flan-T5-XXL encoder [11] with a maximum sequence length of 64 tokens, followed by a tokenwise linear projection from 4096 to 512 dimensions to match the transformer embedding size. We apply classifier-free guidance [23] by independently dropping text embeddings with probability 0.15 and the full conditioning vector with probability 0.05 during training. The model is trained with SiD2 loss [25] (sigmoid shift −3) using the Adam optimizer with a learning rate of 2 × 10−4 and a batch size of 16 for 300K iterations on a single GPU. We initialize training from a pretrained checkpoint to accelerate convergence. The training data consists of HeadsUp latents extracted from our multi-view facial capture dataset, paired with automatically generated text captions describing subject appearance attributes. At inference time, we sample from the learned distribution using a DPM solver [49] with 25 denoising steps. When a text prompt is provided, classifier-free

30

E. Ntavelis et al.

guidance steers the generation toward the described attributes. The sampled latent is then decoded by a frozen pretrained HeadsUp decoder into 3D Gaussian parameters (positions, rotations, scales, opacities, and spherical harmonics color coefficients), producing a complete head avatar that can be rendered in real time. Blendshape-driven Latent Animation As demonstrated in the main text, HeadsUp’s latent space supports controllable facial animation driven by blendshape coefficients. The videos on the supplementary webpage demonstrate that our blendshape-driven animation enables fine-grained control (e.g., eye gaze, asymmetric expressions) and identity-preserving expression transfer from reference performances. Here we explain the architecture used for this experiment. Given a neutral-expression latent and a target blendshape vector, our rigging network predicts a residual that is added to the neutral latent to produce the target expression. Each blendshape value is independently embedded using Fourier features (4 frequency bands), concatenated with a 32-dimensional learnable identifier to distinguish blendshape indices, and projected via a two-layer MLP. The resulting tokens serve as keys and values for an 8-layer, 8-head cross-attention transformer (hidden dimension 1024). The queries are formed by flattening the neutral latent into 1024 tokens. The transformer’s output is then reshaped back to the spatial dimensions of the latent and added to the original neutral representation. We train the latent animation network end-to-end by passing the predicted latent through a frozen pretrained decoder. Supervision is provided by a groundtruth target-expression latent extracted by the encoder. The training objective is a combination of an L1 loss directly on the predicted latent, L1 and multiscale LPIPS perceptual losses on rendered images, and L1 losses on the decoded 3D Gaussian attributes (positions, colors, opacities, rotations, and scales). To preserve high-frequency details, particularly in the eyes and mouth, we employ region-specific LPIPS losses on camera-aware crops and a hinge-style adversarial loss (with a perceptual discriminator) on random 256 × 256 crops, which is activated after 50K steps. The model is trained using Adam with a learning rate of 10−4 and batch size of 80 for 200K steps on 8 A100 GPUs.

F

HeadsUp Training Details

F.1

Dataset processing.

Internal10K Processing. We use our internal multi-view head dataset containing over 10 000 subjects recorded with 16 calibrated RGB cameras. We sample 100 frames per subject for maximum expression diversity. We compute per-view foreground segmentation masks using an internal segmentation model. Ava-256 Processing. Following Avat3r [38], we compute foreground matting masks for the entire dataset using BackgroundMattingV2 [47] and color correct the images to non-linear sRGB.

HeadsUp: Large-scale Gaussian Head Reconstruction

F.2

31

Viewpoint sampling

We select a fixed set of N =10 input cameras from the 16 available views with broad coverage of the face. For models trained or evaluated with fewer than 10 input views (e.g., Avat3r with N =4 or N =6, or ablated HeadsUp models), we use a subset of these 10 cameras selected to maximize face coverage. We prioritize frontal and near-frontal viewpoints before adding side views. Ava-256 For Ava-256, which provides 80 calibrated cameras, we sample N =16 input views via farthest-point sampling on camera positions to ensure maximal viewpoint diversity. Stage 1: Low-Resolution Training. Our model is trained on 2× downsampled images at a resolution of 500 × 375 for 900K steps. We utilize a batch size of 64 and provide 10 input views per training sample. Optimization is performed with Adam [34] with a learning rate of 2 × 10−4 and bfloat16 mixed precision. To ensure stable initialization, we detach the gradients for the opacity and scale parameters during an initial 1K-step warm-up phase. Furthermore, the position regularization weight, λpos , is linearly annealed from 1.0 to 0.01 over the first 100K steps, while the silhouette loss weight, λmask , is also annealed from 2.0 to 0.1 over the same interval. Finally, the adversarial loss, Ladv , is activated at 240K steps. Stage 2: High-Resolution Finetuning. We subsequently continue training at the resolution of 1000 × 750 for 200K steps, using a reduced batch size of 32. During this phase, we introduce region-specific LPIPS perceptual losses applied to eye and mouth crops. To maintain optimization stability, the global LPIPS and discriminator losses are computed on 2× downsampled renders (see Section 3.4 in the main paper). All other loss weights remain identical to those used in Stage 1. We train on the Internal10K dataset until validation metrics plateau using 16 H100 GPUs. A comprehensive summary of all loss weights is provided in Tab. H.3. F.3

Ava-256 Finetuning

Finally, we use the 4 TB version of the Ava-256 dataset [53], which comprises 256 subjects, 80 cameras, and approximately 5000 frames per person. Following the experimental protocol established by Avat3r [38], we train on 244 subjects using 1000 frames per subject, and evaluate our method on 12 held-out validation subjects. We fine-tune the model that was pre-trained on the Internal10K dataset for 200K steps, utilizing 16 input views sampled from the full set of 80 cameras. This fine-tuning stage converges in less than one day using 16 H100 GPUs. F.4

Evaluation Details

Internal10K. We evaluate on 50 held-out validation subjects with 20 frames each, using the same expression-diverse sampling as training. For each frame, we

32

E. Ntavelis et al.

use a fixed set of 10 input views and evaluate on all remaining camera views. All metrics are computed at full resolution (1000 × 750) on composite images (foreground + background). AKD is computed from 2D facial keypoints estimated by PIPNet [28], and CSIM is computed from ArcFace [13] identity embeddings. Ava-256. We evaluate on 12 held-out validation subjects with approximately 2000 sampled frames in total, following Avat3r’s evaluation protocol [38]. For each frame, we use 16 input views selected via farthest-point sampling and evaluate on all remaining cameras. All metrics are computed at full resolution (1024 × 667). AKD and CSIM are computed identically to Internal10K.

G

Potential Negative Societal Impacts

While our framework advances creative industries and telepresence, it presents potential risks regarding the synthesis and manipulation of photorealistic human avatars. High-fidelity 3D head reconstruction lowers the barrier for creating convincing digital humans. Specifically, our blendshape-driven latent animation enables highly controllable rigging of faces into arbitrary expressions. Although our method requires high-quality multi-view studio captures, downstream generative applications could be exploited to generate deepfakes for misinformation, fraud, or non-consensual harassment. To mitigate these risks, we advocate for robust watermarking of synthetic media and emphasize that an individual’s digital likeness must only be created with explicit consent. Regarding data ethics, our model is trained on over 10 000 subjects, all of whom provided written informed consent and received financial compensation. Personally identifiable information is securely managed in compliance with data protection regulations, and all individuals depicted herein explicitly consented to image reproduction. To protect privacy, we verify via nearest-neighbor face similarity that generated subjects do not replicate any training identities. Finally, we acknowledge that demographic imbalances in training data may cause asymmetrical reconstruction quality, disproportionately affecting underrepresented groups. We encourage dataset curation to mitigate such biases in future work.

H

Hyperparameters

For reproducibility, we list important hyperparameters in Tab. H.3. Table H.3. Hyperparameters. Hyperparameter Input and Output Image resolution (Internal) Image resolution (Ava-256)

Value 1000 × 750 1024 × 667 Continued on next page

HeadsUp: Large-scale Gaussian Head Reconstruction Hyperparameter Input views (Internal) Input views (Ava-256) Foreground Gaussians Background Gaussians

33 Value 10 16 65K 262K

Feature Extraction Patch size Patch embedding dimension d Foreground feature dimension c UV latent resolution hf ×wf Background feature dimension c′

7×7 256 512 64 × 64 512

Cross-Attention Transformer Transformer blocks Attention heads Hidden dimension dz MLP dimension Latent resolution hz ×wz

8 8 512 1024 64 × 64

Decoder Number of upsampling stages Residual blocks per upsampling stage Upsampling Block Channel Dimensions Convolutional Kernel Size Activation Functions Position offsets Opacities Scales Rotations Colors (SH coefficients)

2 2 [512, 256] 3×3 tanh ·δmax sigmoid exponential L2 normalization identity

Optimization Optimizer Learning rate Precision Stage 1 iterations Stage 1 batch size Stage 2 iterations Stage 2 batch size

Adam 2 × 10−4 bfloat16 900K 64 200K 32

Losses Warm-up iterations (opacity / scale) Adversarial loss activation iteration Discriminator crop size SH degree L Position offset bound δmax (fg) Position offset bound δmax (bg) λL1

1 000 240K 256 × 256 1 200 mm 10 mm 1.0 Continued on next page

34 Hyperparameter λLPIPS λadv λpos λmask λTV

E. Ntavelis et al. Value 0.1 0.25 1.0 → 0.01 (linearly annealed over 100 K steps) 2.0 → 0.1 (linearly annealed over 100 K steps) 10.0

Record · ID 155271 · SHA-256 45dd2b2cad9944d0
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.