Sony Brings Fair Human-Centric Image Benchmark to Mozilla Data Collective
Sony’s Fair Human-Centric Image Benchmark (FHIBE) joins Mozilla Data Collective, expanding access to globally diverse,
Press Release Disclaimer: This is a press release distributed through the XPR Media network. It has not been independently verified by our newsroom.
![]()
Mozilla Data Collective, the data-sharing platform for human agency and fair value exchange, today announced that Sony’s Fair Human-Centric Image Benchmark (FHIBE) will now be hosted through its platform. FHIBE is a first-of-its-kind globally diverse, consensually collected dataset designed to help AI researchers and developers evaluate fairness, bias, and model utility across human-centric computer vision tasks.
First released by Sony in November 2025, FHIBE includes more than 10,000 images of nearly 2,000 people across 81 countries and regions. The dataset combines detailed annotations and metadata for evaluating model performance across different populations with a consent-based approach to data collection, rather than relying on images scraped from the web without the knowledge of the people pictured.
Bringing FHIBE to Mozilla Data Collective builds on a shared belief between Sony and Mozilla Data Collective that respecting the people behind AI data shouldn’t come at the expense of building useful technology. Mozilla Data Collective gives people and organisations more agency over how their data is shared and used, while helping AI builders access ethical, curated and diverse datasets. Sony’s approach to FHIBE puts those same principles into practice: technically useful data, collected with consent and designed to help technology work more fairly across the people it is meant to serve.
“Consent and human agency shouldn’t be treated as barriers to building useful AI. They should be part of how we build it,” said E.M. Lewis-Jong, Founder and CEO of Mozilla Data Collective. “Sony deserves enormous credit for what they’ve built with FHIBE. It’s a technically rigorous resource that gives developers a meaningful way to evaluate their models, built with the rights of the people represented in the data considered from the start. Making FHIBE available through Mozilla Data Collective expands the kind of high-quality, responsibly sourced data builders can find on the platform. We’re excited to do more to diversify the data available for computer vision, just as we have with speech.”
Sony chose Mozilla Data Collective for its commitment to informed consent and its ability to support GDPR-compliant consent revocation. Sony Group Corporation will continue to own FHIBE, while Mozilla Data Collective will host the dataset for download and manage consent revocations going forward. Sony is also adding approximately 500 images to FHIBE as part of the move. The dataset’s existing terms, conditions, and use cases remain unchanged.
“Our goal with creating FHIBE was to raise the standards for ethical data curation for AI, with informed consent and fair compensation being key principles,” said Alice Xiang Global Head of AI Governance and Lead Research Scientist, Sony Group Corporation. “We are excited to partner with Mozilla Data Collective given these shared commitments and hope that hosting FHIBE on their platform will further expand its reach to the broader community of developers committed to responsible data practices.”
Expanding What’s Available to AI Builders
The addition of FHIBE comes as Mozilla Data Collective continues to expand the range of datasets available through its platform. Thousands of vetted, curated datasets across 450+ languages and cultures are now available, spanning speech, text, image, video and multimodal data.
Sony joins a growing group of organisations choosing Mozilla Data Collective to share high-quality data with AI builders. With FHIBE, that now includes a rigorous evaluation and benchmark dataset designed specifically to assess fairness, bias, and model utility in computer vision models. Every organisation and dataset on the platform is vetted, helping remove some of the guesswork around whether and how data can be used.
Mozilla Data Collective is expanding both what’s available on the platform and what data providers and AI builders can do with it. In July, it launched Compensated Datasets, allowing verified data providers to set their own pricing and licensing terms while receiving 100 percent of the licence fee they set. Participating organisations include Pangeanic, TAUS, Karya, Spotlite, ContentX Labs and YUX Design. Its Lost in Transcription competition challenges developers to build better speech recognition for underserved, code-switching language communities using datasets available through Mozilla Data Collective, while new tools are making it easier for builders to find and request the data they need.
Together, this work advances Mozilla Data Collective’s mission to build a more multilingual, multicultural and multimodal technology future, while rebuilding a data economy where the people and organisations behind datasets have real agency over how their data is used and valued.
FHIBE is available through Mozilla Data Collective here.
About Mozilla Data Collective
Mozilla Data Collective is a mission-locked British social enterprise, backed and incubated by Mozilla Foundation, building the data platform for human agency and fair value exchange. Mozilla Data Collective enables communities, organisations, and individuals to share global cultural datasets on their own terms, while helping downloaders build more representative and culturally grounded technologies with data they cannot find anywhere else. Built by the team behind Mozilla’s Common Voice, the world’s largest open, public-participation speech dataset, Mozilla Data Collective already supports more than 350 organisations sharing over 1,700 datasets across more than 450 languages. Learn more at mozilladatacollective.com.
View source version on businesswire.com: https://www.businesswire.com/news/home/20261008294825/en/
Media gallery