---
title: FACS-based Blendshape Representation
url: https://www.emergentmind.com/topics/facs-based-blendshape-representation
type: topic
---

# FACS-based Blendshape Representation

Searching arXiv for the cited papers and closely related work on FACS-based blendshape representations.
FACS-based blendshape representation is a facial parameterization in which a neutral mesh is combined with a set of semantically meaningful deformation channels aligned with Facial Action Coding System (FACS) action units or closely related animation controls. In its classical form, the face is written as a neutral shape plus weighted blendshape deltas, so that individual channels correspond to localized actions such as jaw open, lip corner pull, brow raise, blink, or cheek motion. Across recent systems, this representation appears in several concrete instantiations rather than as a single universal basis: 52 ARKit-compatible channels for monocular mobile tracking, 53 ICT-FaceKit expression blendshapes for mesh-agnostic cloning and inverse rendering, 32 AU-specific basis vectors in AUBlendSet, 55 FACS-based blendshapes in ICT-FaceKit-based talking-head synthesis, and up to 155 shapes in fully automatic rigging pipelines [2309.05782] [2505.22416] [2507.12001] [2507.20452] [2606.08043].

## 1. Core formulation and control semantics

The canonical mathematical form is the linear blendshape model. One explicit formulation uses a neutral mesh \(\mathbf{b_0}\), fully activated blendshape meshes \(\mathbf{b_i}\), and scalar activations \(w_i\in[0,1]\):
$$
\mathbf{b} = \mathbf{b_0} + \sum_{i=1}^{52} w_i (\mathbf{b_i} - \mathbf{b_0}).
$$
A closely related form appears in automatic rigging systems as
$$
\mathbf{v}(\mathbf{w}) = \mathbf{v}_0 + \sum_{i=1}^N w_i \,\Delta \mathbf{v}_i,
$$
with \(\mathbf{v}_0\) the neutral mesh, \(\Delta \mathbf{v}_i\) the per-shape displacement field, and \(N\) the number of channels [2309.05782] [2606.08043].

The defining property of the FACS-based variant is not merely linearity but semantic factorization. Each channel is intended to encode an interpretable local movement rather than an arbitrary PCA mode. In the ARKit-compatible 52-channel setting, this includes jaw and mouth movements, cheek and nose actions, eyelid and eye squint or widen, brow raise and lower, and lip roll, press, and stretch; the representation is explicitly described as strongly inspired by FACS-style action units [2309.05782]. In AUBlendSet, the formulation becomes even more direct: one AU corresponds to one blendshape basis vector, and a continuous AU coefficient vector \(\mathbf{a}\) synthesizes an expression as
$$
\mathbf{M}_{\text{expr}} = \mathbf{M}_0 + \sum_{i=1}^{N} a_i \mathbf{B}_i,
$$
with \(N=32\) retained AUs on FLAME topology [2507.12001].

A recurrent distinction in the literature is between semantically authored FACS-style controls and data-driven expression bases. Several works explicitly retain artist-designed or AU-indexed channels because they are easier to constrain, easier to retarget, and more interpretable for animation than latent coordinates. One formulation states this directly: animators want to adjust “smile” rather than “mode 3” [2309.05782]. By contrast, statistical slider systems such as SliderGAN use sparse-PCA expression components and continuous parameters in \([-1,1]\); these are blendshape sliders, but not explicit FACS channels [1908.09638].

## 2. Relation to FACS, ARKit, and AU taxonomies

A frequent misconception is that “FACS-based” denotes a single fixed basis. The cited systems instead define several semantically related control spaces, with different cardinalities and different degrees of explicit AU correspondence.

| System | Control space | Relation to FACS |
|---|---:|---|
| Blendshapes GHUM | 52 blendshapes | ARKit-compatible, strongly FACS-inspired |
| ICT-FaceKit-based systems | 53 or 55 blendshapes | FACS-style expression controls |
| AUBlendSet / AUBlendNet | 32 AUs | One AU = one blendshape basis vector |
| OmniFaceRig | up to 155 shapes | Canonical FACS set with side-specific variants and correctives |
| RigAnyFace | 48 primary FACS poses + 48 correctives | Industry-standard FACS pose library |

The 52-channel ARKit-compatible design in Blendshapes GHUM is explicitly chosen because it is familiar to 3D studios and animators, while remaining close to FACS semantics such as AU12-like smiling, AU1-like inner brow raise, AU26-like jaw drop, and AU5-like upper lid raise [2309.05782]. AUBlendSet makes the AU correspondence explicit by starting from the standard FACS AU set, removing AUs related to head pose and tongue, and merging some highly similar eye AUs, yielding 32 retained units; each character then receives 32 AU-Blendshape basis vectors on a 5,023-vertex FLAME mesh [2507.12001].

ICT-FaceKit occupies an intermediate position. In mesh-agnostic cloning and talking-head synthesis, it supplies 53 or 55 expression blendshapes whose labels are FACS-like and semantically localized, including eye, brow, cheek, jaw, and mouth controls [2505.22416] [2507.20452]. In JOLT3D, these 55 blendshapes are the main expression parameterization for both reconstruction and lip-sync; 35 of them are designated as mouth-related and can be replaced independently for audio-driven modification [2507.20452].

At the large-rig end, OmniFaceRig organizes its output into Core, Additional, and Full configurations: approximately 13 basic controls, approximately 46 higher-resolution controls, and a Full set of 155 shapes with side-specific variants and correctives. The paper does not publish a full AU-to-shape table, but it explicitly describes the result as a FACS-based rig and situates it in the space of production-style brow, eyelid, nose, mouth, jaw, viseme-like, and corrective controls [2606.08043]. RigAnyFace similarly targets 48 primary FACS poses and 48 corrective poses as artist-authored blendshapes [2511.18601].

## 3. Acquisition, fitting, and personalization of FACS-style blendshapes

One major line of work obtains FACS-style coefficients by fitting a semantic rig to registered 3D data. Blendshapes GHUM uses a lab setup with calibrated multi-camera capture at 60 Hz, scanning 6,000 identities, each performing 40 predefined expressions chosen so that every blendshape is activated in at least one clip. Scans are registered to a canonical 12,201-vertex template using 478 facial landmarks and an as-conformal-as-possible procedure; a technical artist’s 52 canonical blendshapes are then transferred to each subject’s neutral mesh by affine deformation transfer, and per-frame coefficients are optimized with L-BFGS together with a global rigid transform. The result is a 52-dimensional coefficient vector with ARKit/FACS semantics for every frame, obtained without manual coefficient annotation [2309.05782].

A second line of work personalizes a fixed FACS-style library from minimal input. “Dynamic Facial Asset and Rig Generation from a Single Scan” assumes a generic template rig of 55 additive blendshape displacement fields whose naming convention follows Apple ARKit and includes additional asymmetric eyebrow shapes. It constructs 26 predefined FACS expressions as binary combinations of these 55 units, then learns subject-specific offsets \(\Delta S_i^j\) from a single neutral scan so that personalized blendshapes satisfy
$$
S_i^j = \Delta S_i^j + S_i,
$$
and subject expressions are reconstructed by
$$
P_k^{j}{}' = S_0^j + \sum_{i=1}^{N} \alpha_{ik}^j S_i^j.
$$
A two-stage self-supervised scheme first estimates personalized offsets under locality-aware regularization and then refines both offsets and blending weights to account for unintended motions in real FACS performances [2010.00560].

AUBlendSet provides a more explicit AU library. It contains 500 identities, each with a neutral template mesh and 32 AU-Blendshape basis vectors authored by professional artists using FACS exemplar images and descriptions. The dataset also includes AU-annotated facial postures with continuous intensities rather than only binary or discrete labels. Because all meshes share FLAME topology, the learned basis vectors are directly comparable across characters and can be used to train AU-conditioned generators such as AUBlendNet [2507.12001].

These pipelines share a common structural assumption: identity is stored in the neutral mesh and expression is stored in semantically indexed deformation fields. The personalization mechanisms differ—offline optimization from scans, self-supervised prediction from neutral geometry, or direct artist authoring—but all preserve the channel identity of the original FACS-style template [2309.05782] [2010.00560] [2507.12001].

## 4. Neural, anatomical, and implicit reformulations

Recent work increasingly preserves the FACS-style control interface while changing the internal deformation model. In “Anatomically Constrained Implicit Face Models,” the public interface remains a jaw transform, expression weights \(\mathbf{w}^* \in \mathbb{R}^{N-1}\), and optional head transform, but the neutral surface and per-expression correctives are represented by SIREN-based implicit fields rather than stored vertex deltas. The deformed skin is written as
$$
\mathbf{s}^* = \mathbf{T}_g^*\Big(\text{LBS}\big(\widetilde{\mathbf{s}_0,\mathbf{T}_b^*,\widetilde{k}\big) + \sum_{i=1}^{N-1} w_i^*\,\mathbf{B}_{e}[i]\Big),
$$
so the model remains formally analogous to a jaw-rigged blendshape rig while adding dense anatomical constraints derived from a learned internal anatomy surface [2312.07538].

“High-Quality Mesh Blendshape Generation from Face Videos via Neural Inverse Rendering” retains a standard linear expression model,
$$
\mathbf{v}(\boldsymbol{\beta}_{\text{id}}, \boldsymbol{\beta}_{\exp}) = \mathbf{b}_\mu + B_{\text{id}} \boldsymbol{\beta}_{\text{id}} + B_{\exp} \boldsymbol{\beta}_{\exp},
$$
with \(M_{\exp}=53\) semantically labeled ICT expression blendshapes, but reparameterizes per-vertex deformation through differential coordinates over a tetrahedral mesh. Locality, sparsity, and symmetry regularizers are used specifically to stop optimized personalized bases from losing AU-like semantics during joint inverse rendering of geometry, appearance, and motion [2401.08398].

“Neural Face Skinning for Mesh-agnostic Facial Expression Cloning” makes the FACS supervision explicit at the latent-code level. It uses ICT-FaceKit with 53 expression blendshapes and supervises the first 53 dimensions of a 128-dimensional global expression code \(z_{GE}\) to match the ICT expression weights:
$$
L_{\text{Exp}} = \bigl\| z_{GE}^{\text{exp}} - z_{GT}^{\text{exp}}\bigr\|_2^2 + \bigl\| z_{GE}^{\text{ext}} \bigr\|_2^2.
$$
A per-vertex skinning encoder then localizes this global FACS-interpretable code by predicting vertex-wise weights \(\omega_{\text{Skin}}^{(i)}\) and forming localized codes
$$
z_{LE}^{(i)} = \omega_{\text{Skin}}^{(i)} \odot z_{GE},
$$
which drive a decoder on arbitrary target meshes, including stylized ones [2505.22416].

A different reformulation is compression rather than replacement. “Compressed Skinning for Facial Blendshapes” treats the input rig—often a large FACS-derived blendshape system—as a black box and bakes it into proxy-bone linear blend skinning. Starting from the classical delta form
$$
\hat{v}_i(c)=\hat{v}_{0,i}+\sum_{k=1}^{S} c_k(\hat{v}_{k,i}-\hat{v}_{0,i}),
$$
it learns sparse skinning weights and sparse per-blendshape transform coefficients so that runtime evaluation can use a small number of affine bone transforms instead of hundreds of dense shapes. The FACS control semantics remain unchanged; only the geometric realization is compressed [2406.11597].

## 5. Retargeting, stylization, talking heads, and speech-driven control

Because the channels are semantically labeled, FACS-based blendshape representations are especially suited to cross-identity control and modality transfer. AUBlendNet predicts, from a neutral identity mesh alone, the full set of 32 AU-Blendshape basis vectors for that character. It does this by decoding a style-aware latent code through a learned AUCodebook and then applying the standard linear AU synthesis rule. On AUBlendSet, the reported errors for single-AU and multi-AU control are \(4.55 \times 10^{-9}\) and \(7.62 \times 10^{-9}\), respectively, and the resulting rigs are used for stylized expression manipulation, speech-driven emotional animation, and AU-recognition data augmentation [2507.12001].

JOLT3D uses the 55 ICT-FaceKit FACS-based blendshapes as the expression subspace of a jointly trained reconstruction-and-generation system for talking heads. ReconNet predicts identity \(\alpha\), blendshape coefficients \(\beta\), head pose, and eye pose from RGB frames; a feature-warping GAN then conditions on rendered sketches and 3DMM-induced flow fields. The same FACS-based \(\beta\) is also the target space for an audio-to-expression diffusion model that outputs 35 mouth-related blendshape coefficients. Because the blendshape representation is semantically localized, JOLT3D can replace only the mouth subset of \(\beta\) during lip-sync and preserve the non-mouth channels, which is used to decouple the original chin contour from the lip-synced chin contour and reduce flickering near the mouth [2507.20452].

Work on speech-expression disentanglement reaches a similar structural decomposition even when it does not explicitly use FACS labels. “Learning Disentangled Speech- and Expression-Driven Blendshapes for 3D Talking Face Animation” models deformation as
$$
\Delta V = B_A\,\mathbf{w}_A + B_E\,\mathbf{w}_E,
$$
with a sparsity loss on the cross-domain coefficients \(\boldsymbol{\epsilon}_{A_0}\) and \(\boldsymbol{\epsilon}_{E_0}\). The paper does not define AUs, but it explicitly notes that the learned expression blendshapes are conceptually close to a FACS-style rig, particularly as a compositional upper-face and affective expression layer that can be added to a speech-driven mouth layer. This suggests a direct route to AU-conditioned talking-face systems by replacing or supervising \(B_E\) with a FACS-aligned basis [2510.25234].

Retargeting across topology is also an important theme. Mesh-agnostic neural face skinning localizes a shared FACS-based global code on arbitrary target meshes by indirect supervision from anatomical segmentation labels, while RigAnyFace and OmniFaceRig generate full production-style FACS rigs automatically for previously unrigged assets [2505.22416] [2511.18601] [2606.08043].

## 6. Automation, evaluation, and open limitations

Fully automatic rig generation extends FACS-based representation beyond captured human heads. OmniFaceRig converts a static surface-only 3D character mesh into an inner-mouth-aware FACS rig with up to 155 blendshapes, procedurally fitted teeth, gums, and tongue, repacked UVs, and collision-aware transfer across humans, humanoids, long-muzzled animals, and short-muzzled animals. Its Full set reaches approximately 155 shapes, and Omni-Bench provides 1,000 generated characters with complete FACS rigs and inner-mouth geometry. On screened Omni-Bench inputs, the reported metrics include MAE around \(0.85\)–\(0.92\) mm, Q95 around \(2.5\)–\(2.7\) mm, penetration rate around \(0.05\)–\(0.08\%\), and success rate up to 99% on screened human and humanoid assets [2606.08043].

RigAnyFace similarly targets scalable neural auto-rigging to industry-standard FACS poses. It learns to deform arbitrary neutral meshes into 48 primary FACS poses plus 48 corrective poses, including assets with multiple disconnected components such as eyeballs. A key contribution is 2D supervision on unrigged neutral meshes, which augments a much smaller set of artist-rigged 3D heads and improves generalization to varied topologies while preserving a classical linear blendshape output compatible with existing DCC tools [2511.18601].

Evaluation criteria in this literature reflect both semantics and geometry. Blendshapes GHUM reports Mean Normalized Error for landmarks, with the real-time blendshape model’s reconstructed landmarks at 3.88% and fully activated canonical blendshape differences around 8.45% MNE, emphasizing projected landmark similarity rather than coefficient accuracy alone [2309.05782]. AIM reports average 3D fitting error over 819 frames and 5 actors, with 0.312 mm for AIM versus 0.095 mm for ALM and much larger errors for global and patch blendshape baselines, framing the trade-off between anatomical realism, speed, and strict fidelity [2312.07538]. AUBlendNet, JOLT3D, and the mesh-agnostic cloning literature supplement geometric errors with user studies, inverse-rigging metrics, lip-sync metrics, and per-segment evaluations [2507.12001] [2507.20452] [2505.22416].

Several limitations recur. One is that semantic alignment does not guarantee a universal basis: some systems are actor-specific, as in AIM, while others are tied to a particular template such as ARKit, ICT-FaceKit, or FLAME [2312.07538] [2309.05782] [2505.22416]. A second is that many pipelines are only FACS-inspired rather than explicitly AU-supervised; SliderGAN, Blendshapes GHUM, and some inverse-rendering systems preserve semantic locality without enforcing a one-to-one AU ontology [1908.09638] [2309.05782] [2401.08398]. A third is that the linear blendshape model, even when enhanced with correctives, dynamic textures, or compression, remains an approximation of nonlinear muscle-skin interactions. This is why anatomical simulation, implicit anatomy fields, and collision-aware inner-mouth modeling continue to appear as complementary directions rather than replacements for FACS semantics [1812.02836] [2312.07538] [2606.08043].

Taken together, the recent literature treats FACS-based blendshape representation less as a single file format than as a stable semantic interface. The neutral-plus-deltas formulation remains central, but it now serves as the control layer for annotation-free monocular tracking, self-supervised personalization from a single scan, inverse-rendered rig reconstruction from video, anatomically constrained implicit models, mesh-agnostic neural retargeting, talking-head synthesis, automatic multi-species rigging, and compressed runtime deployment [2309.05782] [2010.00560] [2401.08398] [2505.22416] [2507.20452] [2606.08043].

Source: https://www.emergentmind.com/topics/facs-based-blendshape-representation