After Meta Connect 2022, one Horizon Worlds selfie turned Meta Avatars into a punchline: "soulless," "dead-eyed," "basically a Wii Mii." Leadership promised better graphics. But "better graphics" is a dangerously vague north star, and the only thing worse than a bad avatar was a "better" one everyone hated.
The catch: there was no baseline. No shared definition of quality for the Store, no scores to track, perception that swung by market (some had zero adoption), and leadership asking for proof, not anecdotes.
Meta Avatars 1.0 · 2022 — the version the internet mocked
Lead researcher, Avatars Store. I:
- synthesized years of past research, app reviews, and competitor products before fielding anything new
- turned the org's 5-dimension Experiential Quality Rubric (EQR) into Store-specific, measurable signals with Data Science
- designed and ran a 3-phase global mixed-methods benchmark
- translated the results into roadmap priorities and an instrument teams re-field every six months
Fun isn't optional in an identity product
Teams treated Fun as a nice-to-have: it's a Store, a utility. I disagreed. Creating an avatar isn't a checkout flow, it's an identity experience. Instead of arguing, I kept Fun in the instrument and tested whether it predicted overall quality. It did — one of the strongest relationships in the benchmark.
Measured, not cut. APAC research already pointed to self-expression as a core engagement driver.
6 markets, 3 platforms. A NORAM-only metric would have misled a global roadmap.
Built to re-field from day one. A snapshot tells you what to fix; a benchmark proves you fixed it.
- How do we make "quality" measurable and actionable for the Avatars Store?
- Where is quality lowest — by dimension, market, and platform?
- Which dimensions drive perceived quality, and which issues hurt users most?
- Define: analytics audit and heuristic evaluation with Data Science, comparative interviews (Meta vs. Roblox, Zepeto), and cross-cultural synthesis from Korea and Japan, so we didn't import e-commerce assumptions into an identity product.
- Benchmark: global survey, n = 1,800 across the US, UK, Brazil, Germany, India, and South Korea, on Facebook, Instagram, and VR. Likert per EQR dimension; ANOVA by country with Tukey HSD.
- Explain: 20 severity interviews sampled from the lowest-scoring segments, to learn why scores were low and how much it hurt.
The survey showed where quality broke. The interviews showed why: users weren't shopping, they were exploring who they are.
- Performance is expectation, not speed. Mobile users rated it lower than VR users despite faster loads. VR users expected immersion to take time.
- Inclusivity is a content problem. Beyond skin tones and body types, people couldn't find hairstyles, clothing, and accessories from their own culture.
- Fun is the product. Trying on looks and seeing themselves come to life was the best part of the experience.
PrincipleCreating an avatar is an identity experience, not a checkout flow.
- Became the org's shared language for quality: adopted across Avatar teams and re-measured every six months to prove improvement, not just claim it
- 4 roadmap shifts: cut perceived waiting on mobile and invested in delight instead of chasing raw VR performance; rebuilt try-on as the product; expanded regionally relevant clothing and accessories; built discovery for exploration with personalized recommendations
- Turned inclusivity from a subjective debate into a quantified, prioritizable quality deficit
- As quality perception improved, so did engagement, adoption, and willingness to spend on self-expression
One self, four releases · 2021 → 2024
The reception · 2024 — "Meta Avatars come to life"
Metrics described directionally; exact figures held per NDA.
- Defining the metric when none exists
- Turned an abstract rubric into signals stakeholders could see in the data
- Mixed-methods sequencing
- Quant for how much and where; qual for why and how badly
- Challenging assumptions with evidence
- Tested "Fun is optional" instead of debating it
- Global, cross-platform rigor
- 6 markets, 3 platforms, ANOVA by country
- Research that outlives the study
- A repeatable benchmark, re-fielded every six months