News Report Technology
March 16, 2023

OpenAI Announces Evals, An Open-Source Software Framework for Evaluating AI Models

In Brief

OpenAI hopes to crowdsource benchmarks for evaluating AI models like GPT-4.

Payment processing company, Stripe, has already used Evals to measure the accuracy of their GPT-powered documentation tool.

OpenAI will be granting GPT-4 access for a limited time to those who contribute high quality evals.

OpenAI Announces Evals, An Open-Source Software Framework for Evaluating AI Models

Alongside the announcement of GPT-4, OpenAI has announced the open-source software framework OpenAI Evals. This tool is designed to create and run benchmarks that evaluate the performance of models like GPT-4. With Evals, OpenAI hopes to crowdsource benchmarks for AI model testing. 

“We use Evals to guide development of our models (both identifying shortcomings and preventing regressions), and our users can apply it for tracking performance across model versions (which will now be coming out regularly) and evolving product integrations,” the company explains in a blog post.

Stripe, a popular payment processing company, has already used Evals to complement its human evaluations and measure the accuracy of their GPT-powered documentation tool.

Developers can use Evals to create and run evaluations that:

  • Use datasets to generate prompts,
  • Measure the quality of completions provided by an OpenAI model, and
  • Compare performance across different datasets and models.

With the open-source code, developers can also write and add a custom Eval as well as several templates that may accommodate different benchmarks. The company has included templates that have been most useful internally, including a template for “model-graded evals,” which GPT-4 can use to check its own work. As an example to follow, the company has created a logic puzzles eval containing ten prompts where GPT-4 fails.

Evals is also compatible with implementing existing benchmarks, including several notebooks implementing academic benchmarks and a few variations of integrating small subsets of CoQA.

While developers will not be paid for contributing Evals, OpenAI will be granting GPT-4 access for a limited time to those who contribute “high-quality evals.” 

The announcement of Evals comes after OpenAI recently said it would stop using data submitted by customers via its API to train or improve its models unless the customers decide to opt in. The company joins Meta in crowdsourcing benchmarks as the latter tasks humans with “finding adversarial examples that fool current state-of-the-art models” for its DynaBench platform.

Read more:

Tags:

Disclaimer

In line with the Trust Project guidelines, please note that the information provided on this page is not intended to be and should not be interpreted as legal, tax, investment, financial, or any other form of advice. It is important to only invest what you can afford to lose and to seek independent financial advice if you have any doubts. For further information, we suggest referring to the terms and conditions as well as the help and support pages provided by the issuer or advertiser. MetaversePost is committed to accurate, unbiased reporting, but market conditions are subject to change without notice.

About The Author

Cindy is a journalist at Metaverse Post, covering topics related to web3, NFT, metaverse and AI, with a focus on interviews with Web3 industry players. She has spoken to over 30 C-level execs and counting, bringing their valuable insights to readers. Originally from Singapore, Cindy is now based in Tbilisi, Georgia. She holds a Bachelor's degree in Communications & Media Studies from the University of South Australia and has a decade of experience in journalism and writing. Get in touch with her via [email protected] with press pitches, announcements and interview opportunities.

More articles
Cindy Tan
Cindy Tan

Cindy is a journalist at Metaverse Post, covering topics related to web3, NFT, metaverse and AI, with a focus on interviews with Web3 industry players. She has spoken to over 30 C-level execs and counting, bringing their valuable insights to readers. Originally from Singapore, Cindy is now based in Tbilisi, Georgia. She holds a Bachelor's degree in Communications & Media Studies from the University of South Australia and has a decade of experience in journalism and writing. Get in touch with her via [email protected] with press pitches, announcements and interview opportunities.

Hot Stories
Join Our Newsletter.
Latest News

The Calm Before The Solana Storm: What Charts, Whales, And On-Chain Signals Are Saying Now

Solana has demonstrated strong performance, driven by increasing adoption, institutional interest, and key partnerships, while facing potential ...

Know More

Crypto In April 2025: Key Trends, Shifts, And What Comes Next

In April 2025, the crypto space focused on strengthening core infrastructure, with Ethereum preparing for the Pectra ...

Know More
Read More
Read more
Google Launches AI Futures Fund To Support Startups With Access To AI Models, Funding, And Technical Expertise
News Report Technology
Google Launches AI Futures Fund To Support Startups With Access To AI Models, Funding, And Technical Expertise
May 13, 2025
Arcium Unveils Dark Pool Demo On Testnet, Empowering Secure On-Chain Trading
News Report Technology
Arcium Unveils Dark Pool Demo On Testnet, Empowering Secure On-Chain Trading
May 13, 2025
Gate.io Upgrades MemeBox To Alpha, Accelerating Its Web3 Ecosystem Expansion
News Report Technology
Gate.io Upgrades MemeBox To Alpha, Accelerating Its Web3 Ecosystem Expansion
May 13, 2025
Bitget Provides Critical Aid To Families Affected By Earthquake In Myanmar
News Report Technology
Bitget Provides Critical Aid To Families Affected By Earthquake In Myanmar
May 12, 2025