News Report Technology
March 16, 2023

OpenAI Announces Evals, An Open-Source Software Framework for Evaluating AI Models

OpenAI Announces Evals, An Open-Source Software Framework for Evaluating AI Models

Alongside the announcement of GPT-4, OpenAI has announced the open-source software framework OpenAI Evals. This tool is designed to create and run benchmarks that evaluate the performance of models like GPT-4. With Evals, OpenAI hopes to crowdsource benchmarks for AI model testing. 

“We use Evals to guide development of our models (both identifying shortcomings and preventing regressions), and our users can apply it for tracking performance across model versions (which will now be coming out regularly) and evolving product integrations,” the company explains in a blog post.

Stripe, a popular payment processing company, has already used Evals to complement its human evaluations and measure the accuracy of their GPT-powered documentation tool.

Developers can use Evals to create and run evaluations that:

  • Use datasets to generate prompts,
  • Measure the quality of completions provided by an OpenAI model, and
  • Compare performance across different datasets and models.

With the open-source code, developers can also write and add a custom Eval as well as several templates that may accommodate different benchmarks. The company has included templates that have been most useful internally, including a template for “model-graded evals,” which GPT-4 can use to check its own work. As an example to follow, the company has created a logic puzzles eval containing ten prompts where GPT-4 fails.

Evals is also compatible with implementing existing benchmarks, including several notebooks implementing academic benchmarks and a few variations of integrating small subsets of CoQA.

While developers will not be paid for contributing Evals, OpenAI will be granting GPT-4 access for a limited time to those who contribute “high-quality evals.” 

The announcement of Evals comes after OpenAI recently said it would stop using data submitted by customers via its API to train or improve its models unless the customers decide to opt in. The company joins Meta in crowdsourcing benchmarks as the latter tasks humans with “finding adversarial examples that fool current state-of-the-art models” for its DynaBench platform.

Read more:

Tags:

Disclaimer

In line with the Trust Project guidelines, please note that the information provided on this page is not intended to be and should not be interpreted as legal, tax, investment, financial, or any other form of advice. It is important to only invest what you can afford to lose and to seek independent financial advice if you have any doubts. For further information, we suggest referring to the terms and conditions as well as the help and support pages provided by the issuer or advertiser. MetaversePost is committed to accurate, unbiased reporting, but market conditions are subject to change without notice.

About The Author

Cindy is a journalist at Metaverse Post, covering topics related to web3, NFT, metaverse and AI, with a focus on interviews with Web3 industry players. She has spoken to over 30 C-level execs and counting, bringing their valuable insights to readers. Originally from Singapore, Cindy is now based in Tbilisi, Georgia. She holds a Bachelor's degree in Communications & Media Studies from the University of South Australia and has a decade of experience in journalism and writing. Get in touch with her via [email protected] with press pitches, announcements and interview opportunities.

More articles
Cindy Tan
Cindy Tan

Cindy is a journalist at Metaverse Post, covering topics related to web3, NFT, metaverse and AI, with a focus on interviews with Web3 industry players. She has spoken to over 30 C-level execs and counting, bringing their valuable insights to readers. Originally from Singapore, Cindy is now based in Tbilisi, Georgia. She holds a Bachelor's degree in Communications & Media Studies from the University of South Australia and has a decade of experience in journalism and writing. Get in touch with her via [email protected] with press pitches, announcements and interview opportunities.

Hot Stories

Top 10 Crypto Games to Invest in 2024

by Viktoriia Palchik
February 27, 2024
Join Our Newsletter.
Latest News

Top 10 Crypto Games to Invest in 2024

by Viktoriia Palchik
February 27, 2024

How to Automate Mining With AI?

by Viktoriia Palchik
February 27, 2024

NFTs & Mining: A Digital Synergy

The rise in usage of the non-fungible tokens has changed the way we see and engage with ...

Know More

AI in Crypto

Explore the ever-evolving realm of artificial intelligence within the cryptocurrency sphere. Discover the transformative impact of AI ...

Know More
Join Our Innovative Tech Community
Read More
Read more
Anonymous Dogwifhat (WIF) Trader Bags $1.4M Profit Investing $310 as the Memecoin Surges 50%
Markets News Report
Anonymous Dogwifhat (WIF) Trader Bags $1.4M Profit Investing $310 as the Memecoin Surges 50%
February 27, 2024
HTX Re-Applies for Hong Kong Virtual Asset Trading License Days After Withdrawal
Markets News Report
HTX Re-Applies for Hong Kong Virtual Asset Trading License Days After Withdrawal
February 27, 2024
How to Automate Mining With AI?
News Report
How to Automate Mining With AI?
February 27, 2024
Craig Wright’s Former Lawyers Challenge Authenticity of Emails Shared by his Wife in COPA Trial
Business News Report
Craig Wright’s Former Lawyers Challenge Authenticity of Emails Shared by his Wife in COPA Trial
February 27, 2024
What You
Need to Know

Subscribe To Our Newsletter.
Daily search marketing tidbits for savvy pros.