Home/ Work/ Meta · Avatars Store
Meta · Avatars Store · Avatars 2.0

What Does Quality Even Mean Here?

Measuring what didn't have a measure — what "better" means when there's no benchmark, no precedent, and no agreement. Constructing, evaluating and benchmarking quality metrics for a billion digital identities.

My RoleLead researcher, Avatars Store team
MethodsCompetitive research, global benchmark surveys, severity interviews
OutcomeA shared definition of what to measure — informing quality initiatives and improving the Store over time through benchmarking
1B
Digital identities in scope
1,800
Users benchmarked on 5 EQR dimensions · 6 markets
Mixed methods
Analytics audit → global survey → 20 severity interviews
01The Challenge

After Meta Connect 2022, one Horizon Worlds selfie turned Meta Avatars into a punchline: "soulless," "dead-eyed," "basically a Wii Mii." Leadership promised better graphics. But "better graphics" is a dangerously vague north star, and the only thing worse than a bad avatar was a "better" one everyone hated.

The catch: there was no baseline. No shared definition of quality for the Store, no scores to track, perception that swung by market (some had zero adoption), and leadership asking for proof, not anecdotes.

Meta Avatars 1.0 in 2022: four legless, low-fidelity avatars floating in a park

Meta Avatars 1.0 · 2022 — the version the internet mocked

02My Role

Lead researcher, Avatars Store. I:

  • synthesized years of past research, app reviews, and competitor products before fielding anything new
  • turned the org's 5-dimension Experiential Quality Rubric (EQR) into Store-specific, measurable signals with Data Science
  • designed and ran a 3-phase global mixed-methods benchmark
  • translated the results into roadmap priorities and an instrument teams re-field every six months
03What I Pushed Back On

Fun isn't optional in an identity product

Teams treated Fun as a nice-to-have: it's a Store, a utility. I disagreed. Creating an avatar isn't a checkout flow, it's an identity experience. Instead of arguing, I kept Fun in the instrument and tested whether it predicted overall quality. It did — one of the strongest relationships in the benchmark.

Fun

Measured, not cut. APAC research already pointed to self-expression as a core engagement driver.

Global

6 markets, 3 platforms. A NORAM-only metric would have misled a global roadmap.

Repeatable

Built to re-field from day one. A snapshot tells you what to fix; a benchmark proves you fixed it.

04Research Questions
  1. How do we make "quality" measurable and actionable for the Avatars Store?
  2. Where is quality lowest — by dimension, market, and platform?
  3. Which dimensions drive perceived quality, and which issues hurt users most?
05Approach
  • Define: analytics audit and heuristic evaluation with Data Science, comparative interviews (Meta vs. Roblox, Zepeto), and cross-cultural synthesis from Korea and Japan, so we didn't import e-commerce assumptions into an identity product.
  • Benchmark: global survey, n = 1,800 across the US, UK, Brazil, Germany, India, and South Korea, on Facebook, Instagram, and VR. Likert per EQR dimension; ANOVA by country with Tukey HSD.
  • Explain: 20 severity interviews sampled from the lowest-scoring segments, to learn why scores were low and how much it hurt.
PerformanceUsabilityTrust & InclusivityDesignFun
06Key Insight

The survey showed where quality broke. The interviews showed why: users weren't shopping, they were exploring who they are.

  • Performance is expectation, not speed. Mobile users rated it lower than VR users despite faster loads. VR users expected immersion to take time.
  • Inclusivity is a content problem. Beyond skin tones and body types, people couldn't find hairstyles, clothing, and accessories from their own culture.
  • Fun is the product. Trying on looks and seeing themselves come to life was the best part of the experience.

PrincipleCreating an avatar is an identity experience, not a checkout flow.

07Impact
  • Became the org's shared language for quality: adopted across Avatar teams and re-measured every six months to prove improvement, not just claim it
  • 4 roadmap shifts: cut perceived waiting on mobile and invested in delight instead of chasing raw VR performance; rebuilt try-on as the product; expanded regionally relevant clothing and accessories; built discovery for exploration with personalized recommendations
  • Turned inclusivity from a subjective debate into a quantified, prioritizable quality deficit
  • As quality perception improved, so did engagement, adoption, and willingness to spend on self-expression
The same avatar across four releases from 2021 to 2024, gaining fidelity and expressiveness

One self, four releases · 2021 → 2024

Press coverage of Meta Avatars 2.0: Meta Avatars come to life

The reception · 2024 — "Meta Avatars come to life"

Metrics described directionally; exact figures held per NDA.

08What This Shows
Defining the metric when none exists
Turned an abstract rubric into signals stakeholders could see in the data
Mixed-methods sequencing
Quant for how much and where; qual for why and how badly
Challenging assumptions with evidence
Tested "Fun is optional" instead of debating it
Global, cross-platform rigor
6 markets, 3 platforms, ANOVA by country
Research that outlives the study
A repeatable benchmark, re-fielded every six months

Next Study

Optym · RouteMAX →

Get in touch

Let's think something through together.