FaceMap: Distortion-Driven Perceptual Facial Saliency Maps

*Meta Reality Labs – first author, 1Reality Labs Research, Sausalito CA & Redmond WA
* Equal contribution shuffling? This work: first author. SIGGRAPH Asia 2024 Conference Paper #39 (TOG 9:4)
FaceMap wide hero – single female head plus distortion sensitivity map eyes to cheeks

Wide hero 1600×900 (16:9) – single female head left + distortion heatmap right. Front-page thumb remains featured.jpg 800×900 single-subject.

FaceMap learns where humans notice distortion and reallocates polys / texels / splats there – SROCC 0.82.

Rotating face compare – FaceMap vs uniform (5s loop)

Front → +45° → Front @65K Gaussians – suppl Fig. 14-15 – wide hero keeps thumb clean

Featured thumb preserved: featured.jpg 800×900 portrait single female head – Wowchemy list & front page single-subject.

16:9 wide 1600×900 ~406KB hero – no stretching of portrait thumb (featured.jpg kept separate).

Abstract

First distortion-driven perceptual metric for faces. Generic mesh saliency measures curvature, not tolerance. When a head is compressed to 5% triangles, 1K Gaussians, or 32² texture, where does quality collapse? We capture human preferences via large-scale 2AFC Thurstonian scaling on 10 identities × 5 degradations × 3 views, decoupled into 64×64 overlapping patches (~48K pairs). ANOVA shows allocation method dominates identity (p=7.1e-65 remesh, 1.5e-21 GS) – face perception generalises. We fit a UV-space UNet (512² in → 256² saliency) anchored on 8 semantic UV points, randomised validation r=0.83 RMSE 0.242 JOD vs SD 0.209. Result: SROCC 0.82 / PLCC 0.79 vs Song'14 0.306/0.234, Nehmé'23 0.19/0.23 (weak per Schober). Eyes > wrinkles > mouth > nostrils > silhouette > cheeks cold, but identity-modulates. At 1% tris we beat uniform 98.6% pref, 91.8% at 4%, 75.4% at 16% → mobile sweet-spot. GS 1K: uniform blurs pupils, ours crisp. Texture quadtree saves ~40% leaves.

Faces have a dedicated fusiform area – uniform LOD destroys eyes leaving teeth intact. FaceMap asks "where does distortion become noticeable?" not "where is interesting". Industrial LODs for codec avatars need 5% geo / 128² textures – FaceMap provides the multiplier.

Taxonomy – 10 Bases × 5 Distortions

10 high-quality scanned heads (5 female /5 male, balanced ethnicity/age, ~30K tris face-only, 4K×4K albedo, 200K Gaussians ref) – adapted from Meta Realistic Head collection.

FamilyTypeLevelsMechanism & Prod. analogue
GeometryMesh quantization6 (30%→5% edge keep)Quadric error + uniform; simulates runtime LOD / Draco quant
GeometryLaplacian smoothing6 λ=0.05→0.5Simulates low-LOD blur / skinning linear artifacts
TextureJPEG / Basis compressed6 QF 5→90Texel blockiness – streaming compression
TextureLow-res mip256→32 downsampleBlurriness – texture streaming LOD
SplatsGaussian sparsity5 262K→1K3DGS decimation – mobile splat budget
Stimuli 10 bases x distortion levels wide hero
Suppl Fig.14 – 10 bases × distortion levels wide hero composition (1600×900) – left head, right patches allocation eyes>wrinkles>mouth>cheeks. Single-subject featured.jpg remains 800×900.

Psychophysics – Distort → Render → Patch → 2AFC → JOD

Stimuli & Patching

  • 3 views front 0°, left 45°, right 45° – studio HDRI + rim, 65cm 30° FoV 120 nits sRGB 2.2 D65 calibrated.
  • Overlapping 64×64 patches stride 32 – ~180 per view ~540 per condition – decouples head size & background.
  • Participants see patches only vs reference patch, never full head during forced choice.

2AFC Thurstone

Which patch better vs reference? 200ms ISI, unlimited time, anchored slider "bad–excellent" per Madhusudana'21:

  • Main: 10×5×6×3 = 900 base trials per participant via adaptive QUEST 100 subset.
  • Remesh valid.: 4×6×3×3+12 = 228 trials avg 40 min.
  • GS valid.: 10×5×2×2 = 100 trials avg 34 min (Fig.12/14).
  • N=45+ main 18–42y normal/corrected, 2 outliers >3 MAD removed, gamma-corrected display 2560×1440 65ppd.

JOD: 1 JOD = 75% pref in 2AFC = 0.675σ logistic – Fig.13 Thumb.

FaceMap pipeline method diagram – 1200x1160 padded
Pipeline: distort mesh/tex/splats → 3-view render → patchify → crowd 2AFC → JOD → N-way ANOVA → UNet saliency. 842×814 original upscaled to 1200×1160 with 5% white padding to match hero width – no crop, readable labels.

Learning – Semantic Anchors → UNet Saliency

8 UV Landmarks Anchor

We define 8 semantic UV anchors (eye corners L/R, nose tip, mouth corners, chin bottom, forehead center) seed=6 box. Positions barycentrically interpolated to mean shape UV 512×512. Randomized validation Fig.20: picking 8 random points → Pearson r=0.83 with original, ρ=0.74 RMSE 0.242 JOD vs bootstrap SD 0.209 – robust, not overfit to anchor choice.

Model

  • Input: 512² UV PE + mean curvature + albedo luminance
  • Arch: 4-level UNet 32→256 ch GroupNorm, predicts 256² saliency map
  • Loss: L2 vs empirical JOD + TV + symmetry bilateral regulariser
  • Training: Adam 1e-3 200ep 10-fold leave-one-identity-out CV

Accuracy 10-fold: SROCC 0.82 PLCC 0.79 RMSE 0.31 JOD

SROCC scatter FaceMap 0.82 vs Song 0.306 – 1600x900 readable
Fig.21 Correlation – 1600×900 scatter, axes 0.0–1.0 SROCC/PLCC, tight diagonal FaceMap predicted loss SROCC 0.82 / PLCC 0.79 vs Song'14 0.306/0.234 & Nehmé'23 0.19/0.234 weak per Schober 2018 – labels 12pt preserved for readability.
Qualitative heat: eyes > eye wrinkles / crow's feet > mouth interior > nostrils > silhouette > cheeks/forehead cold. Yet cold spots identity-modulated – e.g., freckled cheeks slightly warmer, bearded chin moderate sensitivity.

Applications – Polys / Texels / Splats Reallocation

Applications allocation polys texels splats – 2-row stacked 1400x1308
Applications: remesh / texture / GS reallocation – eyes & mouth get 2–3× budget vs uniform. 1400×1308 stacked composite displayed 100% width with rounded corners & shadow ().

Remesh / LOD Allocation

Weighted quadric weight = FaceMap(x)·curv(x)^0.5.
BudgetPref vs Uniform
1% tris ultra-low98.6%
4%91.8%
16%75.4%
65%54.1% n.s.
Mobile sweet-spot – perceptual priors matter when bandwidth-limited.

Gaussian Splatting 3DGS

Allocate counts per facial region ∝ FaceMap. Fig.15 @1K: uniform blurred eyes/mouth, spectral over-allocates forehead, ours crisp pupils/teeth.
Same anchor interpolation works across UV connectivities.

Texture Quadtree Compression

Non-salient leaves → mean color, saving ~40% leaves same perceptual SSIM. Suppl A.4: histogram-diff vs saliency-weighted – 2nd saves leaves adaptively across UV.
Works across different UV charts thanks to anchor design.

Video & Interactive

Placeholder – replace with SIG Asia archive when released. Suppl includes HTML hover viewer (WebGL diff uniform vs FaceMap).

Links & BibTeX

Last rebuilt Aug 13 2026 from suppl parsing. Photo credit Ria Zhang. Single-female-head thumb kept per spec.

@inproceedings{jiang2024facemap,
  title={FaceMap: Distortion-Driven Perceptual Facial Saliency Maps},
  author={Jiang, Zhongshi and Venkateshan, Kishore and Nam, Giljoo and Chen, Meixu and Bachy, Romain and Bazin, Jean-Charles and Chapiro, Alexandre},
  booktitle={SIGGRAPH Asia 2024 Conference Papers},
  number={39},
  pages={1--11},
  year={2024},
  publisher={ACM},
  volume={9},
  doi={10.1145/3680528.3687631},
  url={https://dl.acm.org/doi/10.1145/3680528.3687631},
  note={Suppl: https://achapiro.github.io/Jia24/Jia24sup.pdf}
}

# Evaluation baseline citations:
@inproceedings{song2014mesh-saliency,
  title={Mesh saliency},
  author={Song, ...}
}
@inproceedings{nehme2023geolpips,
  title={Graphics-LPIPS...}
}

Website template based on the Nerfies project page. If you reuse their source code, please credit them appropriately.