By using this site, you agree to the Privacy Policy and Terms of Use.
Accept

Your #1 guide to start a business and grow it the right way…

InSmartBudget

  • Home
  • Startups
  • Start A Business
    • Business Plans
    • Branding
    • Business Ideas
    • Business Models
    • Fundraising
  • Growing a Business
  • Funding
  • More
    • Tax Preparation
    • Leadership
    • Marketing
Subscribe
Aa
InSmartBudgetInSmartBudget
  • Startups
  • Start A Business
  • Growing a Business
  • Funding
  • Leadership
  • Marketing
  • Tax Preparation
Search
  • Home
  • Startups
  • Start A Business
    • Business Plans
    • Branding
    • Business Ideas
    • Business Models
    • Fundraising
  • Growing a Business
  • Funding
  • More
    • Tax Preparation
    • Leadership
    • Marketing
Made by ThemeRuby using the Foxiz theme Powered by WordPress
InSmartBudget > Startups > This Tool Probes Frontier AI Models for Lapses in Intelligence

This Tool Probes Frontier AI Models for Lapses in Intelligence

News Room By News Room April 3, 2025 5 Min Read
Share

Executives at artificial intelligence companies may like to tell us that AGI is almost here, but the latest models still need some additional tutoring to help them be as clever as they can.

Scale AI, a company that’s played a key role in helping frontier AI firms build advanced models, has developed a platform that can automatically test a model across thousands of benchmarks and tasks, pinpoint weaknesses, and flag additional training data that ought to help enhance their skills. Scale, of course, will supply the data required.

Scale rose to prominence providing human labor for training and testing advanced AI models. Large language models (LLMs) are trained on oodles of text scraped from books, the web, and other sources. Turning these models into helpful, coherent, and well-mannered chatbots requires additional “post training” in the form of humans who provide feedback on a model’s output.

Scale supplies workers who are expert on probing models for problems and limitations. The new tool, called Scale Evaluation, automates some of this work using Scale’s own machine learning algorithms.

“Within the big labs, there are all these haphazard ways of tracking some of the model weaknesses,” says Daniel Berrios, head of product for Scale Evaluation. The new tool “is a way for [model makers] to go through results and slice and dice them to understand where a model is not performing well,” Berrios says, “then use that to target the data campaigns for improvement.”

Berrios says that several frontier AI model companies are using the tool already. He says that most are using it to improve the reasoning capabilities of their best models. AI reasoning involves a model trying to break a problem into constituent parts in order to solve it more effectively. The approach relies heavily on post-training from users to determine whether the model has solved a problem correctly.

In one instance, Berrios says, Scale Evaluation revealed that a model’s reasoning skills fell off when it was fed non-English prompts. “While [the model’s] general purpose reasoning capabilities were pretty good and performed well on benchmarks, they tended to degrade quite a bit when the prompts were not in English,” he says. Scale Evolution highlighted the issue and allowed the company to gather additional training data to address it.

Jonathan Frankle, chief AI scientist at Databricks, a company that builds large AI models, says that being able to test one foundation model against another sounds useful in principle. “Anyone who moves the ball forward on evaluation is helping us to build better AI,” Frankle says.

In recent months, Scale has contributed to the development of several new benchmarks designed to push AI models to become smarter, and to more carefully scrutinize how they might misbehave. These include EnigmaEval, MultiChallenge, MASK, and Humanity’s Last Exam.

Scale says it is becoming more challenging to measure improvements in AI models, however, as they get better at acing existing tests. The company says its new tool offers a more comprehensive picture by combining many different benchmarks and can be used to devise custom tests of a model’s abilities, like probing its reasoning in different languages. Scale’s own AI can take a given problem and generate more examples, allowing for a more comprehensive test of a model’s skills.

The company’s new tool may also inform efforts to standardize testing AI models for misbehavior. Some researchers say that a lack of standardization means that some model jailbreaks go undisclosed.

In February, the US National Institute of Standards and Technologies announced that Scale would help it develop methodologies for testing models to ensure they are safe and trustworthy.

What kinds of errors have you spotted in the outputs of generative AI tools? What do you think are models’ biggest blind spots? Let us know by emailing [email protected] or by commenting below.

Read the full article here

News Room April 3, 2025 April 3, 2025
Share This Article
Facebook Twitter Copy Link Print
Previous Article Ooni Pizza Oven: From Backyard Side Hustle to $200 Million
Next Article Being Sick for 7 Days Exposed a Hard Truth About My Business
Leave a comment Leave a comment

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Wake up with our popular morning roundup of the day's top startup and business stories

Stay Updated

Get the latest headlines, discounts for the military community, and guides to maximizing your benefits
Subscribe

Top Picks

Marketing Online Can Be Overwhelming For Small Businesses — But It Doesn’t Have to Be With These 6 Strategies
July 5, 2025
Why I Think More Startups Should Try Rotating Leadership
July 5, 2025
eBay and Vestiaire Collective Want an Exemption from Trump’s Tariffs
July 5, 2025
Former Marine Turns Health Scare Into B2B Wellness Media Startup
July 5, 2025
Netflix teams up with Yahoo DSP as it builds out ads tier
July 5, 2025

You Might Also Like

eBay and Vestiaire Collective Want an Exemption from Trump’s Tariffs

Startups

Venice Braces for Jeff Bezos and Lauren Sanchez’s Wedding

Startups

Cloudflare Is Blocking AI Crawlers by Default

Startups

Substack Is Having a Moment—Again. But Time Is Running Out

Startups

© 2023 InSmartBudget. All Rights Reserved.

Helpful Links

  • Privacy Policy
  • Terms of use
  • Press Release
  • Advertise
  • Contact

Resources

  • Start A Business
  • Funding
  • Growing a Business
  • Leadership
  • Marketing

Popuplar

The $3.1 Trillion in Value Companies Still Overlook
Sisters’ Side Hustle Leads to Hundreds of Millions of Dollars
Venice Braces for Jeff Bezos and Lauren Sanchez’s Wedding

We provide daily business and startup news, benefits information, and how to grow your small business, follow us now to get the news that matters to you.

Welcome Back!

Sign in to your account

Lost your password?