News Report Technology
March 16, 2023

OpenAI Announces Evals, An Open-Source Software Framework for Evaluating AI Models

In Brief

OpenAI hopes to crowdsource benchmarks for evaluating AI models like GPT-4.

Payment processing company, Stripe, has already used Evals to measure the accuracy of their GPT-powered documentation tool.

OpenAI will be granting GPT-4 access for a limited time to those who contribute high quality evals.

OpenAI Announces Evals, An Open-Source Software Framework for Evaluating AI Models

Alongside the announcement of GPT-4, OpenAI has announced the open-source software framework OpenAI Evals. This tool is designed to create and run benchmarks that evaluate the performance of models like GPT-4. With Evals, OpenAI hopes to crowdsource benchmarks for AI model testing. 

“We use Evals to guide development of our models (both identifying shortcomings and preventing regressions), and our users can apply it for tracking performance across model versions (which will now be coming out regularly) and evolving product integrations,” the company explains in a blog post.

Stripe, a popular payment processing company, has already used Evals to complement its human evaluations and measure the accuracy of their GPT-powered documentation tool.

Developers can use Evals to create and run evaluations that:

  • Use datasets to generate prompts,
  • Measure the quality of completions provided by an OpenAI model, and
  • Compare performance across different datasets and models.

With the open-source code, developers can also write and add a custom Eval as well as several templates that may accommodate different benchmarks. The company has included templates that have been most useful internally, including a template for “model-graded evals,” which GPT-4 can use to check its own work. As an example to follow, the company has created a logic puzzles eval containing ten prompts where GPT-4 fails.

Evals is also compatible with implementing existing benchmarks, including several notebooks implementing academic benchmarks and a few variations of integrating small subsets of CoQA.

While developers will not be paid for contributing Evals, OpenAI will be granting GPT-4 access for a limited time to those who contribute “high-quality evals.” 

The announcement of Evals comes after OpenAI recently said it would stop using data submitted by customers via its API to train or improve its models unless the customers decide to opt in. The company joins Meta in crowdsourcing benchmarks as the latter tasks humans with “finding adversarial examples that fool current state-of-the-art models” for its DynaBench platform.

Read more:

Tags:

Disclaimer

In line with the Trust Project guidelines, please note that the information provided on this page is not intended to be and should not be interpreted as legal, tax, investment, financial, or any other form of advice. It is important to only invest what you can afford to lose and to seek independent financial advice if you have any doubts. For further information, we suggest referring to the terms and conditions as well as the help and support pages provided by the issuer or advertiser. MetaversePost is committed to accurate, unbiased reporting, but market conditions are subject to change without notice.

About The Author

Cindy is a journalist at Metaverse Post, covering topics related to web3, NFT, metaverse and AI, with a focus on interviews with Web3 industry players. She has spoken to over 30 C-level execs and counting, bringing their valuable insights to readers. Originally from Singapore, Cindy is now based in Tbilisi, Georgia. She holds a Bachelor's degree in Communications & Media Studies from the University of South Australia and has a decade of experience in journalism and writing. Get in touch with her via [email protected] with press pitches, announcements and interview opportunities.

More articles
Cindy Tan
Cindy Tan

Cindy is a journalist at Metaverse Post, covering topics related to web3, NFT, metaverse and AI, with a focus on interviews with Web3 industry players. She has spoken to over 30 C-level execs and counting, bringing their valuable insights to readers. Originally from Singapore, Cindy is now based in Tbilisi, Georgia. She holds a Bachelor's degree in Communications & Media Studies from the University of South Australia and has a decade of experience in journalism and writing. Get in touch with her via [email protected] with press pitches, announcements and interview opportunities.

The DOGE Frenzy: Analysing Dogecoin’s (DOGE) Recent Surge in Value

The cryptocurrency industry is rapidly expanding, and meme coins are preparing for a significant upswing. Dogecoin (DOGE), ...

Know More

The Evolution of AI-Generated Content in the Metaverse

The emergence of generative AI content is one of the most fascinating developments inside the virtual environment ...

Know More
Join Our Innovative Tech Community
Read More
Read more
Arbitrum Foundation Proposes Expansion Program Adjustment To Enable Deployment Of New Orbit Chains Across Networks Beyond Ethereum
News Report Technology
Arbitrum Foundation Proposes Expansion Program Adjustment To Enable Deployment Of New Orbit Chains Across Networks Beyond Ethereum
April 18, 2024
Blast’s DEX Thruster Finance Raises $7.5M In Funding From Pantera Capital And OKX Ventures To Enhance On-Chain Experience For Users
Business News Report Technology
Blast’s DEX Thruster Finance Raises $7.5M In Funding From Pantera Capital And OKX Ventures To Enhance On-Chain Experience For Users
April 18, 2024
State of DePIN 2024 Report Reveals Key Insights From Decentralized Physical Infrastructure Networks Landscape
Markets News Report
State of DePIN 2024 Report Reveals Key Insights From Decentralized Physical Infrastructure Networks Landscape
April 18, 2024
Solana-Based Derivatives Protocol Zeta Markets Unveils Tokenomics, Allocates 10% Of Token Supply For Airdrops
Markets News Report Technology
Solana-Based Derivatives Protocol Zeta Markets Unveils Tokenomics, Allocates 10% Of Token Supply For Airdrops
April 18, 2024