Before the Shutter: Aesthetic and Actionable Portrait Photography Planning in 3D Scenes RUIXIANG JIANG, The Hong Kong Polytechnic University, China CHANG WEN CHEN, The Hong Kong Polytechnic University, China Inputs
Aesthetic Portrait Planning (before the shutter)
Captured Portrait
arXiv:2605.30318v1 [cs.GR] 28 May 2026
Human subject
Prompts
“Thoughtful” “Pianist” “Graceful”
Interactive planning
creative?
… and more dramatic?
Selected Plan 2
Can u make it more …
Fig. 1. Given a 3D scene, a human subject, and user prompts, our system generates candidate portrait plans before capture by jointly exploring subject pose, camera placement, lighting, and exposure in 3D space. Each plan is visualized in the scene and linked to its rendered viewfinder image. Iterative prompting and comparative feedback refine the plan frontier toward the desired photographic intent.
Portrait photography is largely decided before the shutter opens: the subject’s pose, the camera configuration, and the lighting devices must be coordinated within the surrounding 3D scene. In contrast, most existing computational methods focus on post-production in 2D image space, such as retouching, relighting, or editing images that already exist; pre-capture photographic planning remains largely unexplored. We introduce 3D aesthetic portrait planning, the task of generating human pose, camera, lighting, and exposure plans that produce visually compelling portraits while satisfying geometric and photometric feasibility in a 3D scene. Our approach builds a Photographic Scene Graph that represents scene affordances, subject-scene relations, and portrait-relevant lighting structure. Built on this representation, we perform aesthetic-guided comparative planning over previous attempts and current viewfinder observations. Experiments across diverse indoor and outdoor scenes show that our method produces portraits preferred by human raters and MLLM evaluators over competitive baselines, while maintaining high physical plausibility. Together, our results suggest a path from
post-capture correction toward pre-capture computational portrait planning. Project repository: https://github.com/songrise/Before-the-Shutter.
1
Introduction “You don’t take a photograph, you make it.” Ansel Adams
Portrait photography is one of the most universal and enduring forms of visual expression. From professional studio sessions to casual smartphone snapshots, capturing a person in a way that is readable and appealing remains a widely shared aspiration. Advances in camera hardware and computational photography have made cameras more capable than ever: a sharp, well-exposed image is now trivial to achieve with a single click. Yet the problem of
2
•
Jiang and Chen
making an aesthetically compelling portrait, where the pose communicates, the light sculpts, and the composition feels considered, remains as challenging as ever. To understand what makes a portrait, we can separate the photographic process into pre-shutter planning and post-shutter refinement [Adams 1981; Hunter et al. 2015] stages. Given a scene and a human actor, pre-shutter planning organizes the relationship among the subject, the camera, and the lighting to aesthetically communicate the desired effect. The difficulty lies in the fact that these choices are both coupled and scene-dependent: pose, camera, lens, exposure, and lighting each change how the others read, and the same setup can succeed in one environment but fail in another. This abundance of interacting choices is the expressive power of pre-shutter planning: for a trained photographer, pose, viewpoint, lens, and light become instruments for aesthetic intent. At the same time, knowing how to balance these choices is precisely what makes portrait photography difficult. Despite the importance of this pre-shutter planning, most computational approaches to portraiture begin only after an image has been captured. While paradigms such as retouching [Su et al. 2025], relighting [Rao et al. 2024], re-composition through cropping [Zhong et al. 2021], and general image editing [Wan Team 2026] demonstrate their power in 2D refinement, they are structurally confined to 2D image space: they cannot access the 3D degrees of freedom that govern pre-shutter decisions, and therefore cannot produce physically actionable plans. Such reactive approaches may suggest what a more aesthetic portrait could look like, but they do not answer how to make it. For example, a retouching method may brighten a shadowed face, but it does not tell us where to place the key and fill lights, or how to set their relative power, to avoid flat or uncontrolled lighting. A cropping method may improve 2D composition, but it cannot choose the camera position and focal length needed to preserve background context while keeping the subject visually balanced. All of these important pre-shutter planning decisions remain outside the scope of current systems. This leaves a gap between image-space enhancement and the scene-level decisions that photographers actually control before capture. In this paper, we ask whether these pre-shutter decisions can be modeled computationally. Given a static 3D scene, a human subject, and a user prompt specifying the desired portrait style, our goal is to generate a set of candidate portrait plans – each specifying a subject pose and placement, camera configuration, controllable lighting, and exposure – that are both visually compelling and physically actionable in the scene. We call this task 3D aesthetic portrait planning. It is an inverse problem by nature: aesthetic quality is observed only in the resulting 2D viewfinder image, yet the available controls are 3D and tightly coupled. This asymmetry makes feed-forward prompt-to-plan prediction brittle: the aesthetic effect of pose, camera, and light often becomes clear only through scene-grounded interaction and observation. We therefore propose a planning algorithm that simultaneously understands the spatial and photometric structure of the scene, and iteratively refines a frontier of candidate plans through comparative aesthetic judgement. Our contributions are fourfold. First, we formulate aesthetic portrait planning as prompt-conditioned generation of subject pose, camera, lighting, and exposure plans, and introduce a benchmark
of 50 tasks spanning 14 diverse indoor and outdoor scenes. Second, we propose a Photographic Scene Graph that jointly encodes spatial affordances, subject-scene relations, and portrait-relevant lighting structure, enabled grounded portrait planning. Third, we introduce Aesthetic-Guided Comparative Planning, a trainingfree algorithm that maintains an aesthetic frontier of candidate plans and refines it through comparative aesthetic signal, supporting both MLLM-as-Judge for fully automatic planning and human judgement for interactive refinement. Fourth, we instantiate these components with coordinated Photographer, Actor, and Judge roles and show that it improves human and MLLM aesthetic preference over image-only, one-pass, greedy, and spatial-graph baselines while maintaining high actionability.
2 Related Work 2.1 Computational portrait aesthetics. Computational portrait methods, including retouching [Chang et al. 2026; Elezabi et al. 2024], relighting [Pandey et al. 2021; Sun et al. 2019], and composition [Yuan et al. 2024], operate primarily in 2D image space on captured images, which makes them post-shutter by design. To improve attractiveness, these methods use Image Aesthetic Assessment (IAA) models [Kirstain et al. 2023; Murray et al. 2012; Schuhmann 2022] as aesthetic proxies. However, canonical IAA score is not interpretable and usually provide limited guidance for converting an image-level aesthetic preference into concrete 3D changes in pose, camera, or lighting. Recent work suggests that MLLMs encode human-aligned and interpretable aesthetic priors than IAA models [Jiang and Chen 2025], motivating pairwise comparison as a more flexible planning signal than absolute scoring. We instead propose a training-free planning approach supporting both MLLM-as-Judge and human judge, where an aesthetic frontier is refined through pairwise comparative judgement. This facilitates translation of image aesthetic signal into actionable plans in 3D scene.
2.2
Scene-aware human placement and interaction.
Human-scene interaction (HSI) methods recover or synthesize plausible human bodies in 3D environments by reasoning over geometry, semantic affordances, contact relations, and physical constraints [Hassan et al. 2019; Li and Dai 2024; Savva et al. 2016; Zhang et al. 2020; Zhao et al. 2022]. These methods ground the body in the surrounding scene – chairs afford sitting, walls afford leaning, contacts constrain stable pose – but are evaluated purely on physical plausibility. This is insufficient for portrait photography in three respects: a physically stable pose may be visually illegible from the camera, poorly lit by available illumination, or unexpressive under the intended prompt. Our work inherits physical grounding from this literature as an abstract Actor role, which is coordinated with the Photographer and Judge to jointly satisfy camera legibility, illumination, and aesthetic preference.
2.3
Camera control and portrait lighting design.
Camera placement and lighting are fundamental creative controls in computer graphics, but they have largely been studied in isolation from each other and from subject pose. Intelligent camera
Before the Shutter: Aesthetic and Actionable Portrait Photography Planning in 3D Scenes
control and virtual cinematography methods plan viewpoints in 3D scenes under constraints such as visibility, framing, shot size, composition, continuity, and aesthetics [Christie et al. 2008; Galvane et al. 2015; Halper and Olivier 2000; Lino et al. 2011; Liu et al. 2024; Xie et al. 2023; Zhang et al. 2025]. These systems show that camera placement is an expressive scene-level decision, but they usually assume that actor pose and illumination are fixed. Portrait relighting and reflectance-capture methods approach the problem from the opposite direction, demonstrating the importance of light direction, softness, and subject-background separation [Mei et al. 2024; Nestmeyer et al. 2020; Pandey et al. 2021], yet most of them still assume that the subject and camera have already been chosen and operate in 2D. The shared limitation across all of these works is that each assumes the other variables are fixed: camera control fixes pose and light; lighting design fixes subject and camera; portrait relighting fixes everything and works in 2D. To our knowledge, prior work has not addressed the joint pre-shutter planning of subject pose, camera configuration, and controllable lighting for promptconditioned portrait photography in 3D scenes. We close this gap through the Photographic Scene Graph, which jointly encodes spatial affordances, scene-human relations, and lighting structure in a unified representation, enabling holistic portrait planning.
3 Methodology 3.1 Preliminaries Scene Representation (𝐸). We consider both the geometry and photometry for the shooting environment and use a unified representation for real-world and virtual scenarios. First, a global occupancy field O : R3 → {0, 1} conceptually encodes rigid scene geometry, where occupied regions are used primarily to validate shooting actionability. For the photometric structure, we require the scene to have a radiance function. A camera ray with origin o and outgoing direction 𝝎 𝑜 then queries the scene radiance, which we denote as 𝐿(o, 𝝎 𝑜 ). In virtual environments, this radiance is produced by physically-based rendering (PBR). Human Representation (H ). We represent the subject with SMPLX. Treating shape, facial expression, and appearance, garments as fixed inputs, we plan the body & hand pose 𝜽 𝑏 , 𝜽 ℎ and root transform (R, t) ∈ 𝑆𝐸 (3), writing the human state as H = {𝜽 𝑏 , 𝜽 ℎ , R, t}. This state induces an articulated mesh used as a geometric proxy for pose planning and body–scene collision checking. Controllable Lighting Representation (L): In portrait photography, scene lighting is often supplemented with controllable lighting equipment, including light sources (e.g., key, fill, and rim lights) and non-emissive occluders such as negative fill. These devices are used to sculpt the subject for aesthetic purposes [Hunter et al. 2015]. Consequently, we decompose the total incident radiance 𝐿𝑖 into the environmental contribution 𝐿𝑖,𝑒𝑛𝑣 and the controllable photographic contribution 𝐿𝑖,𝑐𝑡𝑟𝑙 : 𝐿𝑖 (x, 𝝎 𝑖 ) = 𝐿𝑖,𝑒𝑛𝑣 (x, 𝝎 𝑖 ) + 𝐿𝑖,𝑐𝑡𝑟𝑙 (x, 𝝎 𝑖 ; L)
(1)
We parameterize L = {𝑙 1, . . . , 𝑙 𝑁𝑙 } as controllable lighting devices. Each device 𝑙𝑛𝑙 is defined by a rigid transformation in 𝑆𝐸 (3), and
•
3
geometric shape and size; emitters additionally include radiometric power and color or temperature. Camera Representation (C) and Image Formation: We adopt a thinlens camera model that unifies geometric projection with physicallybased exposure simulation. The camera is parameterized as C = (T, K, P), where T ∈ 𝑆𝐸 (3) defines the camera extrinsic and K encodes focal length 𝑓 and principal point, together determining the ray origin o and direction 𝝎 𝑜 for each image coordinate. We use an aperture-priority camera, with photographic controls P = {𝑁 𝑓 , Δ}, where the f-number 𝑁 𝑓 controls depth-of-field and Δ denotes exposure compensation in stops relative to the scenemetered baseline. ISO is fixed at 100 for notation simplicity, and the shutter time 𝜏 is determined by camera metering; we assume an idealized shutter without motion blur. Under this setting, the exposure value EV100 is calculated as: EV100 = log2
𝑁 𝑓2 .
(2)
𝜏
For each image coordinate u ∈ Ω, the camera model back-projects u into a ray (o, 𝝎 𝑜 (u)); we write L(u) ≡ 𝐿(o, 𝝎 𝑜 (u)) as shorthand for the rendered radiometric signal along that ray. The final image is produced by a shutter imaging operator S, which abstracts the image formed when the planned camera is triggered: I(u) = S(𝐸, H, C, L) u = Γ 𝜙 𝜅 2− (EV100 +Δ) L(u) , (3) where 𝜅 is a calibration constant absorbing lens transmittance, and 𝜙 is a monotone saturating response before the tone-mapping operator Γ, reflecting finite exposure latitude of camera sensors [Chen et al. 2019; Lin et al. 2011]. Overall, this formulation couples geometric composition (T, K) with photographic control (𝑁 𝑓 , Δ) while leaving S as an abstract mapping from scene state to the final image.
3.2
Problem Formulation
We introduce the task of aesthetic portrait photography planning. Given an initially unobserved static scene 𝐸, a human subject A providing fixed identity shape and appearance and rigging, and an optional user prompt 𝑦 specifying the desired portrait style or constraints, our goal is to generate an aesthetic portrait plan 𝑠 ∗ = (H ∗, C ∗, L ∗ ), where H ∗ specifies the subject pose and placement, C ∗ specifies the camera configuration, and L ∗ specifies the controllable lighting setup. The plan should satisfy two key constraints: (1) Actionability. The plan should be executable in the scene: the subject, camera, and lights occupy valid free-space configurations; the human body avoids body-scene penetration; and the pose should be stable without floating. The planned exposure should be feasible given the scene’s photometric structure and the camera’s dynamic range. (2) Photographic Aesthetics. The resulting viewfinder image should be judged preferable to alternative candidate plans under the prompt. This preference is assessed in terms of subject-scene coherence, expressive pose and placement, compositional intent, and lighting design.
4
•
Jiang and Chen
Update
Photographic Scene graph
Aesthetic-Guided Comparative Planning
D>A
staging
B.2
C.2
D
…
key 8:1
fill
composition
3:1 : Emissive
A
B.1
C.1
…
lighting : Non-emissive
A > C.1
Stage: composition
Anchor
Fig. 2. Pipeline Overview. We progressively construct a Photographic Scene Graph to ground aesthetic-guided comparative planning. Left: the graph represents scene nodes (e.g., window, lamp, bookshelf), the human subject, controllable lights, and their spatial and photometric relations. Right: the comparative planning loop, shown with composition as an example, where the Photographer iteratively proposes candidate plan states and the Judge compares each new viewfinder observation against the frontier to accept, refine, or reject the state.
3.3
Aesthetic-Guided Comparative Planning
To coordinate physical actionability with 2D photographic aesthetics, our key insight is to use relative aesthetic judgement to guide planning with decoupled roles. Specifically, our system includes three roles: a Photographer, an Actor, and a Judge. The Photographer propose edits to plan state, including camera, lighting C, L, and staging (i.e., human pose and placement) proposal H . The Actor attempts to realize the proposed staging in the environment subject to physical actionability. Both roles directly manipulate the plan state, resulting in a new viewfinder observation. The Judge does not directly manipulate the plan but instead provides comparative aesthetic judgment on the observations, which drives the planning process. We denote a planning state as 𝑠𝑡 = (H𝑡 , C𝑡 , L𝑡 ), and its resulting viewfinder image as I𝑡 = S(𝐸, 𝑠𝑡 ), where 𝑡 indexes the planning step. During planning, we maintain an aesthetic frontier 𝐾𝑡 F𝑡 = {(𝑠𝑖 , I𝑖 , 𝑎𝑖 )}𝑖=1 , where 𝑠𝑖 = (H𝑖 , C𝑖 , L𝑖 ) is an accepted best-sofar candidate state, I𝑖 = S(𝐸, 𝑠𝑖 ) is its viewfinder observation, and 𝑎𝑖 stores stage and judgement metadata. At each step, the Judge compares the new observation I𝑡 relative to frontier observations {I𝑖 } (𝑠𝑖 ,I𝑖 ,𝑎𝑖 ) ∈ F𝑡 , and decides based on relative aesthetic judgment whether the new state 𝑠𝑡 should be added to the frontier, further refined, or discarded. When refinement is required, the Photographer proposes a revision Δ𝑠 to the current plan state and applies it to transition to the next state 𝑠𝑡 +1 . During staging, this additionally requires the Actor to realize the revision under physical constraints. This process iterates until a stopping criterion is met, such as reaching the maximum planning budget. In this way, we integrate 2D aesthetic preference into a state-space planning process. Comparative aesthetic guidance directs the search toward more promising planning directions, as illustrated in Fig. 2.
3.4
Photographic Scene Graph
Planning to improve the aesthetics in I𝑡 is an inverse problem that necessitates understanding the interactions between H, C, L and the scene 𝐸. To achieve this, we introduce Photographic Scene Graph
(Photographic SG). A Photographic SG G is a semantic representation of 𝐸 and 𝑠𝑡 , recording both spatial affordances and lighting structure of the scene: G = (V𝑛𝑜𝑛 , V𝑒𝑚𝑖 , E𝑛2𝑛 , E𝑒2𝑛 , E𝑒2𝑒 ).
(4)
Non-emissive nodes V𝑛𝑜𝑛 include both scene objects of interest (e.g., a chair or a tourist attraction) and body parts of interest (e.g., face, torso), storing attributes including affordance tags and reflected-light metering EV100 . Emissive nodes V𝑒𝑚𝑖 represent principal portrait-relevant light sources, including dominant environmental emitters and controllable light devices introduced during planning. We define three types of relations: (1) spatial relations between non-emissive nodes E𝑛2𝑛 , (2) directed emissive-to-non-emissive light influence E𝑒2𝑛 , and (3) emissive-to-emissive relative source strength E𝑒2𝑒 . For the E𝑒2𝑒 , we let Photographer actively probe the scene 𝐸 so that it can understand its photometric structure with minimal calibration. Specifically, for each emitter 𝑣 𝐴 ∈ V𝑒𝑚𝑖 , we estimate its isolated contribution to a subject by ambient subtraction, 𝑀𝐴Δ = 𝑀𝐴+amb −𝑀amb , where 𝑀 denotes the spot-metered luminance of a Lambertian probe placed at the subject anchor, and 𝑀amb can be estimated by turning off all controllable lights and occluding dominant environmental emitters. We then record the emitter-toambient ratio and pairwise emitter ratio as: 𝑀𝐴Δ 𝑀𝐴Δ 𝑑𝐴 2 , 𝑣 𝐴 , 𝑣 𝐵 ∈ V𝑒𝑚𝑖 , (5) 𝑟 𝐴:amb = , 𝑟 𝐴:𝐵 = Δ 𝑀amb 𝑀𝐵 𝑑𝐵 where 𝑑𝐴 , 𝑑𝐵 are light-to-subject distances. This normalization reduces dependence on absolute calibration and provides a practical proxy for the concept of lighting ratio in portrait photography [Hunter et al. 2015]. The Photographer progressively constructs and updates the Photographic SG G𝑡 at each planning step using viewfinder observations, camera readings, and active probes of viewpoint and lighting. This construction is semi-automated: the MLLM-based Photographer parses each stepwise observation into semantic nodes and
Before the Shutter: Aesthetic and Actionable Portrait Photography Planning in 3D Scenes
spatial relations, while reflected-light metering and photometric relations are algorithmically probed through active interaction with the 3D scene.
3.5
SG-Anchored Staging, Composition, and Lighting
We now describe the stage-wise planning procedure, where the evolving photographic SG provides shared spatial and photometric context across staging, composition, and lighting. Staging. Staging decides the subject pose and root transform. Fundamentally, we view this problem as harmonizing E𝑛2𝑛 between human and scene nodes by considering affordance and aesthetics. Photographer queries G for affordance-compatible non-emissive nodes, such as seats, and uses them as candidate anchors for placing H . Each placement proposal is grounded in the image plane, tracked across nearby views, and used to condition a 2D generative prior for proposing staging sketches. To lift these 2D proposals into 3D, we estimate SMPL-X pose 𝜽 𝑏 , 𝜽 ℎ from 2D proposals via SMPLer-X [Cai et al. 2023], instantiate the body with the subject-specific shape 𝜷 induced by A, and calculate root transform through multi-view triangulation. The Actor then attempts to realize H in 𝐸 subject to affordance and physical constraints, and the Judge provides imagelevel feedback to drive the next staging proposal. Composition. We use the term composition to denote planning of C. The goal of composition is to organize the spatial relationship among subjects within the view frustum. In portraiture, both the human and scene nodes such as landmarks and furniture can serve as subjects. The Photographer therefore queries G to perform constrained composition planning. This constraint has two aspects: it controls the visibility of desired nodes 𝑣 ∈ V𝑛𝑜𝑛 and guides the balance of inter-subject relationships E𝑛2𝑛 for aesthetic purposes. At each step, the Photographer reviews the viewfinder observation I𝑡 against the two constraints and proposes a revision ΔC to the camera parameters. Lighting. This stage controls the lighting devices and camera exposure. Conditioned on the user prompt and the Judge’s critique, Photographer first selects a portrait lighting pattern such as Rembrandt lighting, as a structured initialization around an anchor node 𝑣 ∈ V𝑛𝑜𝑛 . Specifically, each preset defines a group of light devices with relative position and strength. It then iteratively refines each controllable light’s rigid transformation, size, power, and color temperature, or tries another preset. As in real portrait photography [Hunter et al. 2015], this refinement is essential yet tricky: small changes in source position, size, or power can substantially alter facial modeling, subject separation, shadows, and exposure under the ambient scene light. The Photographic SG guides this process in two ways. First, E𝑒2𝑛 and E𝑒2𝑒 encode how existing emitters illuminate the subject anchor and their relative strength. It informs the selection of devices to complement, reshape, or counteract the ambient contribution. Second, the reflected-light attribute EV100 on 𝑣 ∈ V𝑛𝑜𝑛 functions as a semantic meter for the scene, constraining the Photographer to keep the intended subjects within the camera dynamic range and to match the desired tone.
•
5
4 Experiments and Results 4.1 Implementation Details For the virtual environment, we use Blender with the Cycles renderer. All assets and lighting setups are configured to be physically plausible, including scale, materials [Burley 2012], and light sources. For actor staging, we use Gemini-Flash-2.5-Image (i.e., NanoBanana) for 2D pose proposal, CoTracker [Karaev et al. 2024] for cross-view anchor tracking, and SMPLer-X-H32 [Cai et al. 2023] for SMPL-X body recovery. For the planning system, we implement Photographer as an MLLM agent by prompting a GPT-5.4-mini model, and the Judge as GPT-5.4 by default. For a complete planning round, we empirically set the maximum number of steps to 3 for staging and 7 for composition and lighting. We implement the system on a server with a single NVIDIA RTX 4090-D GPU. Per-task runtime ranges from 5–25 minutes and is bounded by rendering complexity. The detailed system parameters are included in the supplementary materials. The full codebase and benchmark will be released upon publication.
4.2
Evaluation Protocol
Benchmark Set. We construct a benchmark set for aesthetic portrait planning. It contains tasks across 14 static 3D scenes, covering 8 indoor and 6 outdoor environments with diverse lighting conditions. We sample a total of 50 tasks with diverse user prompts for evaluation. For each case, a method receives the scene, the character, and the prompt, then outputs a complete portrait plan with its associated rendering. Baselines. To the best of our knowledge, no existing system addresses the joint planning of human pose, camera configuration, and lighting control in a unified 3D portrait framework. We therefore compare against baselines that represent alternative degrees of photographic reasoning and scene grounding. Random Planner samples valid cameras in free space, places a rest-pose human in front of the camera, and uses preset lighting. Template Photographer samples SMPL poses, with eye-level or three-quarter cameras, and standard portrait lighting as a rule-based lower bound. Image-Only Planner plans iteratively according to viewfinder images and the prompt, without scene graph anchoring. Spatial-Graph Planner uses affordances and spatial relations without photometric nodes and edges for iterative refinement. Photographic-Graph One-Pass uses the full graph to generate a complete plan in one step, without iterative revision. Photographic-Graph Greedy uses the full graph and refines iteratively, but it exploits local view aesthetic signals without using the aesthetic frontier. We compare these baselines against our full coordinated planner.
4.3
Metrics
3D Physical Actionability. We evaluate actionability with metrics: skeletal human-scene collision, static pose balance, and exposure validity. For collision, we report the bone-penetration-free rate 1 ∑︁ 𝑅coll = 1[𝑃skel (𝑠𝑖 ) = 0] , (6) 𝑁 𝑖 where B𝑠 denotes skeletal samples of the SMPL-X bones.
6
•
Jiang and Chen
“Melancholy”
“Mafia”
“Horrifying”
Final planned portrait
Staged scene
“Touching the pagoda”
Fig. 3. Qualitative visualization of our planning approach. Each column shows the input prompt, the staged 3D scene, and the final shoot under planned camera and lighting control. Best viewed in color and zoomed in.
For static balance, we report 𝑅bal =
1 ∑︁ 𝑆 (𝑠𝑖 ), 𝑁 𝑖
𝑆 (𝑠) = 1 𝜋 ⊥𝑔 (CoM H ) ∈ ConvHull(𝜋⊥𝑔 (Q𝑠 )) ,
(7) where 𝜋⊥𝑔 projects points onto the plane perpendicular to gravity, sup Q𝑠 = {x ∈ B𝑠 : 𝑑 O (x) ≤ 𝜖𝑐 } are support-contact samples from load-bearing SMPL-X bones within contact threshold 𝜖𝑐 , and CoM H is estimated using anthropometric segment-mass priors [Winter 2009]. For exposure, we first compute the valid-pixel fraction 𝑝 valid =
i + 1 ∑︁ h −𝑠 − 1 2 ≤ 𝜌 (u) ≤ 2𝑠 , |Ω| u∈Ω
𝜌 (u) =
L(u) , 𝐿mid
(8)
where 𝐿mid is the scene-linear luminance mapped to middle gray under the planned exposure. We report its smoothed logit: 𝑉exp = log
+ 10−6
𝑝 valid , 1 − 𝑝 valid + 10−6
(9)
where Ω is the image plane and 𝑠 − = 𝑠 + = 3 characterize empirical reliable signal-exposure latitude of digital cameras [DXOMARK 2015]. 2D Aesthetic Quality. Generic IAA models [Kirstain et al. 2023; Xu et al. 2023] are not well matched to our task: they judge isolated images or prompt-image preference, and do not explicitly account for whether a portrait uses the given scene, subject, and use prompt effectively under the judgment subjectivity. We therefore use a two-alternative forced choice (2AFC) protocol that assess relatively from 4 dimensions: subject staging, camera composition, lighting & exposure, and overall aesthetic quality. Specifically, we perform round-robin tournament comparisons for photo produced by different baselines under same input, then estimate their performance using Bradley–Terry (BT) models [Bradley and Terry 1952]: Pr(𝑖 ≻ 𝑗) =
exp(𝛽𝑖 ) , exp(𝛽𝑖 ) + exp(𝛽 𝑗 )
(10)
where 𝛽𝑖 and 𝛽 𝑗 denote the latent photographic capabilities of methods 𝑖 and 𝑗, respectively. We report both the point estimate and 95% confidence interval of 𝛽. We collect judgement from both human and an MLLM (Gemini-3pro). Expert raters (n=19) contribute 13140 valid annotations, with inter-annotator agreement (Cohen’s kappa) of 0.66, 0.61, 0.69, 0.73 on each dimension. We also report zero-shot MLLM judgments as a secondary automatic evaluator, motivated by recent findings that MLLMs produces better human-aligned zero-shot aesthetic judgments [Jiang and Chen 2025] under 2AFC setting, compared with IAA models. In our experiment, the MLLM alignment (Spearman’s rho) with human judgment is 0.60, 0.75, 0.92, and 0.78 on each dimension, justifying its use as a complementary evaluator.
4.4
Quantitative Results
Main Results. Tab. 1 reports the final-plan performance across actionability and aesthetic preference. Our full method obtains the highest overall preference under both evaluators, with BT scores of 1.30 ± 0.18 from the MLLM judge and 0.96 ± 0.19 from human raters, while maintaining high physical actionability. SG-based methods show a substantial improvement over the Image-only baselines: the scene graph constrains planning to be scene-aware and helps produce affordance-grounded human poses (higher 𝑅coll , 𝑅bal ) and manageable exposure settings (higher 𝑉exp ). On the other hand, iterative planning is also essential: the Photographic One-pass method, which uses the full Photographic SG but without iterative refinement, performs worse than the iterative-planning-based SpatialGraph Planner. We attribute this to the lack of interaction with the scene for grounding the plan. Comparative planning built on top of the iterative planner substantially improves both actionability and preference scores, as demonstrated by the comparison between our full method and Photographic-Graph Greedy. Stage-wise Ablations. To isolate the contributions of comparative planning (F for short) and the Photographic SG G, we checkpoint the plan at the end of each stage of our full method and fork it into different branches with either module removed. Tab. 2 summarizes
Before the Shutter: Aesthetic and Actionable Portrait Photography Planning in 3D Scenes
•
7
Table 1. Main quantitative results on the portrait planning benchmark. We report collision-free rate, balance, and exposure validity for actionability. Aesthetic columns report Bradley–Terry preference scores with point estimates and 95% confidence intervals. All metrics are higher the better. Actionability Scores
Aesthetics (Bradley-Terry 𝛽)
Method
𝑅coll ↑
𝑅bal ↑
𝑉exp ↑
Staging ↑
Composition ↑
Light & Exposure ↑
Overall (MLLM) ↑
Overall (Human) ↑
Random Planner Template Photographer Image-Only Planner Spatial-Graph Planner Photographic-Graph One-Pass Photographic-Graph Greedy Ours Full
1.00 0.95 0.84 0.94 0.82 0.80 0.92
0.00 0.02 0.18 0.58 0.42 0.52 0.56
1.08 0.93 1.32 1.49 1.41 1.35 1.56
−1.93±0.27 −1.15±0.22 0.47±0.19 0.80±0.19 0.29±0.19 0.59±0.19 0.92±0.19
−1.81±0.24 −0.25±0.17 −0.10±0.17 0.61±0.17 0.19±0.17 0.64±0.17 0.72±0.17
−1.50±0.25 −1.06±0.21 0.23±0.19 0.29±0.20 0.19±0.19 0.61±0.20 1.25±0.23
−2.13±0.25 −1.71±0.22 0.12±0.17 0.84±0.17 0.50±0.17 1.07±0.18 1.30±0.18
−1.43±0.20 −0.72±0.17 0.15±0.16 0.47±0.17 0.01±0.17 0.54±0.17 0.96±0.19
“Solemn”
“Resting on pier” “Waiting for someone”
“Confident”
Ours
Greedy
Spatial Graph
ImageOnly
“Le penseur, dramatic”
Fig. 4. Qualitative comparison of our method and baselines. Compared with ours, baselines less faithfully coordinate pose, camera, and lighting to match the prompt. The generated pose can be floating or awkward, the composition can be unbalanced or even miss the subject(s), and the lighting can be indiscriminately flat, harsh, or underexposed, failing to match the desired tone implied by prompt. Zoom in for details.
the ablation-internal scores using the MLLM evaluator. Adding G allows the planner to make informed adjustments to camera and lighting. However, this does not readily improve planning quality in the absence of F . This is because G and its induced constraints are progressively constructed during planning. We empirically find that planning without F is prone to self-reinforcing state dynamics. This observation is reminiscent of cycling and local lock-in in memoryless local search in traditional control theory literature [Glover 1989; Mayne et al. 2000]. We visualize this phenomenon in Fig. 5. Combining both F and G achieves the best result by performing scene- and state-aware planning.
Table 2. Stage-wise ablation of Comparative Planning (F) and Photographic SG (G). no-op means skip the stage. Higher is better. Method
Staging 𝛽 ↑
Comp. 𝛽 ↑
Light 𝛽 ↑
no-op +F +G +( F, G)
– −0.25±0.27 −0.05±0.26 0.29±0.27
−0.39±0.26 −0.02±0.24 −0.14±0.24 0.55±0.24
−0.31±0.25 −0.27±0.25 0.12±0.25 0.46±0.24
•
Ours Full
w/o Comparative Planning
Ours Full
w/o Comparative Planning
8
Jiang and Chen
Stage: composition
“Close-up of a powerful fist punching directly towards the camera lens, wide-angle.”
1
2
3
4
5
6
1
2
3
4
5
6
Stage: lighting
1
“Resting against the table, Chiaroscuro, Tyndall, moody.” power +100w
split_light(1.5m)
2
3 split_light(1m)
Planned Lighting
EV + 0.5, dist–=0.5m
Power+=300w
Key,620w
...
6 4
6
EV+=0.5; power+=200w Ev-=0.8,Negative_fill()
Ev+=0.3, rim_light()
negativefill
rim,80w Key,300w
... 1
2
3
4
6
Fig. 5. Ablation of Comparative Planning. Comparative planning allows the planner to revert to earlier frontier states when a refinement direction is judged to degrade aesthetics. In the top example, it helps to solve the ambiguity between wide-angle and camera distance. In the bottom example, it informs the planner to use a negative fill on the composite side to enhance the contrast instead of continuously strengthening the key light. Green outline: accepted frontier state.
4.5
Qualitative Results and Visualization
Fig. 3 visualizes the prompt, staged scene, and the final captured portrait for four sample tasks. The proposed method produces aesthetic plans that coordinate the human subject and scene, composed with cinematic camera angles and lighting. Fig. 4 compares the result produced by our method with selected baselines. The results show that our method produces more visually appealing portraits that better harmonize the scene and subject and adhere better to the prompt. Fig. 6 visualizes of how the Photographic SG anchors planning in composition and lighting stage.
4.6
Limitations and Future Work
In this work, the staging is limited to the body and hand pose, while leaving the important facial expression and gaze direction fixed. This omission is due to two reasons: (1) Technically, intricate facial expressions make it challenging to transfer expression between 2D guidance and an arbitrary 3D articulated human asset at high fidelity. (2) Evaluation of facial expression is more challenging. In practice we tested human assets with expressions and find users suffer from uncanny valley effects when rendered realistically, which can bias the judgment [Tinwell et al. 2011]. This is also why we only utilized the mannequin asset in our experiments. We therefore leave facial-expression planning as a future direction, which may require joint learning with an expression prior and use a more specialized evaluation protocol.
5
Conclusion
We introduced aesthetic portrait planning in 3D scenes, a task that shifts computational portraiture from post-capture 2D editing to precapture 3D decision making. Our method combines a Photographic SG with aesthetic-guided comparative planning to jointly search over pose, camera, and lighting while respecting physical feasibility in the scene. This formulation allows us to solve the inverse problem of mapping 2D portrait aesthetics to 3D actionable plans through scene-grounded interaction and observation. Experiments show that the proposed method improves both actionability and human preference over competitive baselines. We hope this setting opens new opportunities for scene-aware photographic assistance, virtual production, and embodied creative planning.
Before the Shutter: Aesthetic and Actionable Portrait Photography Planning in 3D Scenes
Photographic Scene Graph (at stage start, simplified) lit
red_chair
window
Sit-on human_torso
face
ambient
left-of
hand
tables
Stage constraint (generated by MLLM Judge):
0.07
torso EV100: -3.9
face EV100: -2.0
glass_wall upper_body EV100: -0.2
Stage End Viewfinder w/o SG
Stage constraint (generated by MLLM Judge):
Photographic Scene Graph (at stage end, simplified) floor
ambient
Stand-on
wall_light
“Keep face and upper_body, glass_wall within roughly ±2.5 EV latitude. Avoid washing out while preserving the mid-tolow-key tone of the minimalist interior scene. Add a soft, warm key light from high […]
key key
Stand-on torso EV100: 0.8
Stage End Viewfinder
Stage End Viewfinder w/o SG
fill
6:1
fill
0.43 0.03
face
ambient
EV100: 1.0
0.03
glass_wall upper_body
wall_light
EV100: 1.2
EV100: 0.8
EV100: -3.2
Stage Start Viewfinder
9
“Keep face and hand centered in frame to establish the moody pose. Keep red_chair in the view to establish the sit_on relationship”
Photographic Scene Graph (at stage start, simplified) floor
Stage End Viewfinder
•
Planned Lighting (Ours)
Key,165w (#FFECDA)
fill,25w (#FFEEDF)
Fig. 6. Visualization of Photographic Scene Graph. Top: An example of SG-anchored composition, prompt: “Melancholy”. The MLLM judge generated constraint to guide the composition to preserve the human pose and the red chair to establish the mood. Bottom: An example of SG-anchored lighting, prompt: “Gracefully dancing near the glass”. The SG provide photometric structure of the scene, which is used to guide the lighting design and exposure planning. Simplified means we manually select the relevant nodes and edges for visualization readability. Zoom in for details.
10
•
Jiang and Chen
References Ansel Adams. 1981. The Negative. New York Graphic Society, Boston. Ralph Allan Bradley and Milton E. Terry. 1952. Rank Analysis of Incomplete Block Designs: I. The Method of Paired Comparisons. Biometrika 39, 3/4 (1952), 324–345. Brent Burley. 2012. Physically-Based Shading at Disney. In ACM SIGGRAPH 2012 Courses. Zhongang Cai, Wanqi Yin, Ailing Zeng, Chen Wei, Qingping Sun, Wang Yanjun, Hui En Pang, Haiyi Mei, Mingyuan Zhang, Lei Zhang, et al. 2023. Smpler-x: Scaling up expressive human pose and shape estimation. Advances in Neural Information Processing Systems 36 (2023), 11454–11468. Zewei Chang, Zheng-Peng Duan, Jianxing Zhang, Chun-Le Guo, Siyu Liu, Hyungju Chun, Hyunhee Park, Zikun Liu, and Chongyi Li. 2026. PerTouch: VLM-Driven Agent for Personalized and Semantic Image Retouching. Proceedings of the AAAI Conference on Artificial Intelligence 40, 4 (2026), 2752–2759. doi:10.1609/aaai.v40i4. 37264 Can Chen, Scott McCloskey, and Jingyi Yu. 2019. Analyzing Modern Camera Response Functions. In Proceedings of the IEEE Winter Conference on Applications of Computer Vision. 1961–1969. Marc Christie, Patrick Olivier, and Jean-Marie Normand. 2008. Camera Control in Computer Graphics. Computer Graphics Forum 27, 8 (2008), 2197–2218. doi:10.1111/ j.1467-8659.2008.01181.x DXOMARK. 2015. Nikon D7200: The New APS-C Champ. https://www.dxomark.com/ nikon-d7200-the-new-aps-c-champ/ Accessed: 2026-05-03. Omar Elezabi, Marcos V. Conde, Zongwei Wu, and Radu Timofte. 2024. INRetouch: Context Aware Implicit Neural Representation for Photography Retouching. arXiv preprint arXiv:2412.03848 (2024). Quentin Galvane, Rémi Ronfard, Christophe Lino, and Marc Christie. 2015. Continuity Editing for 3D Animation. In Proceedings of the AAAI Conference on Artificial Intelligence. 753–762. doi:10.1609/aaai.v29i1.9288 Fred Glover. 1989. Tabu search—part I. ORSA Journal on computing 1, 3 (1989), 190–206. Nicolas Halper and Patrick Olivier. 2000. CAMPLAN: A Camera Planning Agent. In Smart Graphics: Papers from the 2000 AAAI Spring Symposium. AAAI Press, 92–100. Mohamed Hassan, Vasileios Choutas, Dimitrios Tzionas, and Michael J. Black. 2019. Resolving 3D Human Pose Ambiguities with 3D Scene Constraints. In Proceedings of the IEEE International Conference on Computer Vision. 2282–2292. Fil Hunter, Steven Biver, Paul Fuqua, and Robin Reid. 2015. Light: Science and Magic: An Introduction to Photographic Lighting (5 ed.). Focal Press. Ruixiang Jiang and Changwen Chen. 2025. Multimodal LLMs Can Reason about Aesthetics in Zero-Shot. In Proceedings of the 33rd ACM International Conference on Multimedia. 6634–6643. doi:10.1145/3746027.3754961 Nikita Karaev, Ignacio Rocco, Benjamin Graham, Natalia Neverova, Andrea Vedaldi, and Christian Rupprecht. 2024. Cotracker: It is better to track together. In European conference on computer vision. Springer, 18–35. Yuval Kirstain, Adam Polyak, Uriel Singer, Shahbuland Matiana, Joe Penna, and Omer Levy. 2023. Pick-a-Pic: An Open Dataset of User Preferences for Text-to-Image Generation. In Advances in Neural Information Processing Systems. Lei Li and Angela Dai. 2024. Genzi: Zero-shot 3d human-scene interaction generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 20465–20474. Haiting Lin, Seon Joo Kim, Sabine Süsstrunk, and Michael S. Brown. 2011. Revisiting Radiometric Calibration for Color Computer Vision. In Proceedings of the IEEE International Conference on Computer Vision. 129–136. Christophe Lino, Mathieu Chollet, Marc Christie, and Rémi Ronfard. 2011. Computational Model of Film Editing for Interactive Storytelling. In Interactive Storytelling. Springer, 305–308. doi:10.1007/978-3-642-25289-1_35 Xinhang Liu, Yu-Wing Tai, and Chi-Keung Tang. 2024. ChatCam: Empowering Camera Control through Conversational AI. arXiv preprint arXiv:2409.17331 (2024). David Q Mayne, James B Rawlings, Christopher V Rao, and Pierre OM Scokaert. 2000. Constrained model predictive control: Stability and optimality. Automatica 36, 6 (2000), 789–814. Yiqun Mei, Yu Zeng, He Zhang, Zhixin Shu, Xuaner Zhang, Sai Bi, Jianming Zhang, HyunJoon Jung, and Vishal M. Patel. 2024. Holo-Relighting: Controllable Volumetric Portrait Relighting from a Single Image. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 4263–4272. Naila Murray, Luca Marchesotti, and Florent Perronnin. 2012. AVA: A Large-Scale Database for Aesthetic Visual Analysis. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2408–2415. doi:10.1109/CVPR.2012.6247954 Thomas Nestmeyer, Jean-François Lalonde, Iain Matthews, and Andreas Lehrmann. 2020. Learning Physics-Guided Face Relighting Under Directional Light. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 5124–5133. doi:10.1109/CVPR42600.2020.00517 Rohit Pandey, Sergio Orts-Escolano, Chloe LeGendre, Christian Häne, Sofien Bouaziz, Christoph Rhemann, Paul Debevec, and Sean Fanello. 2021. Total Relighting: Learning to Relight Portraits for Background Replacement. ACM Transactions on Graphics 40, 4, Article 43 (2021), 21 pages. doi:10.1145/3450626.3459872
Pramod Rao, Gereon Fox, Abhimitra Meka, Mallikarjun B R, Fangneng Zhan, Tim Weyrich, Bernd Bickel, Hanspeter Pfister, Wojciech Matusik, Mohamed Elgharib, and Christian Theobalt. 2024. Lite2Relight: 3D-Aware Single Image Portrait Relighting. In ACM SIGGRAPH 2024 Conference Papers. doi:10.1145/3641519.3657470 Manolis Savva, Angel X Chang, Pat Hanrahan, Matthew Fisher, and Matthias Nießner. 2016. PiGraphs: Learning Interaction Snapshots from Observations. ACM Transactions on Graphics 35, 4, Article 139 (2016), 12 pages. doi:10.1145/2897824.2925867 Christoph Schuhmann. 2022. LAION-Aesthetics. https://laion.ai/blog/laion-aesthetics/ Accessed: 2026-05-06. Wanchao Su, Can Wang, Chen Liu, Fangzhou Han, Hongbo Fu, and Jing Liao. 2025. StyleRetoucher: Generalized Portrait Image Retouching With GAN Priors. IEEE Transactions on Visualization and Computer Graphics 31, 9 (2025), 5089–5100. doi:10. 1109/TVCG.2024.3432910 Tiancheng Sun, Jonathan T. Barron, Yun-Ta Tsai, Zexiang Xu, Xueming Yu, Graham Fyffe, Christoph Rhemann, Jay Busch, Paul E. Debevec, and Ravi Ramamoorthi. 2019. Single Image Portrait Relighting. ACM Transactions on Graphics 38, 4, Article 79 (2019). doi:10.1145/3306346.3323008 Angela Tinwell, Mark Grimshaw, Debbie Abdel Nabi, and Andrew Williams. 2011. Facial expression of emotion and perception of the Uncanny Valley in virtual characters. Computers in Human behavior 27, 2 (2011), 741–749. Wan Team. 2026. Wan-Image: Pushing the Boundaries of Generative Visual Intelligence. arXiv preprint arXiv:2604.19858 (2026). David A Winter. 2009. Biomechanics and motor control of human movement. John wiley & sons. Desai Xie, Ping Hu, Xin Sun, Soren Pirk, Jianming Zhang, Radomír Mech, and Arie E Kaufman. 2023. GAIT: Generating aesthetic indoor tours with deep reinforcement learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision. 7409–7419. Jiazheng Xu, Xiao Liu, Yuchen Wu, Yuxuan Tong, Qinkai Li, Ming Ding, Jie Tang, and Yuxiao Dong. 2023. ImageReward: Learning and Evaluating Human Preferences for Text-to-Image Generation. In Advances in Neural Information Processing Systems. Quan Yuan, Leida Li, and Pengfei Chen. 2024. Aesthetic Image Cropping Meets VLP: Enhancing Good While Reducing Bad. Journal of Visual Communication and Image Representation 105 (2024), 104316. doi:10.1016/j.jvcir.2024.104316 Mengchen Zhang, Tong Wu, Jing Tan, Ziwei Liu, Gordon Wetzstein, and Dahua Lin. 2025. GenDoP: Auto-regressive Camera Trajectory Generation as a Director of Photography. arXiv preprint arXiv:2504.07083 (2025). Siwei Zhang, Yan Zhang, Qianli Ma, Michael J. Black, and Siyu Tang. 2020. PLACE: Proximity Learning of Articulation and Contact in 3D Environments. In International Conference on 3D Vision. 642–651. Kaifeng Zhao, Shaofei Wang, Yan Zhang, Thabo Beeler, and Siyu Tang. 2022. Compositional Human-Scene Interaction Synthesis with Semantic Control. In European Conference on Computer Vision. 311–327. Lei Zhong, Feng-Heng Li, Hao-Zhi Huang, Yong Zhang, Shao-Ping Lu, and Jue Wang. 2021. Aesthetic-Guided Outward Image Cropping. ACM Transactions on Graphics 40, 6, Article 211 (2021), 13 pages. doi:10.1145/3478513.3480566