Loading…

Loading grant details…

Active STANDARD GRANT National Science Foundation (US)

SLES: Monitoring, Improving, and Certifying Safe Foundation Models

$8M USD

Funder National Science Foundation (US)
Recipient Organization University of Illinois At Urbana-Champaign
Country United States
Start Date Sep 01, 2024
End Date Aug 31, 2027
Duration 1,094 days
Number of Grantees 2
Roles Principal Investigator; Co-Principal Investigator
Data Source National Science Foundation (US)
Grant ID 2416897
Grant Description

During the past couple of years, Artificial Intelligence (AI) has had a dramatic impact in many areas. Many of the current AI systems are based on Large language models (LLMs), which are computer models that capture information across a variety of topic domains. LLMs have been increasingly deployed in many high-stakes applications including Web search and recommendation, healthcare and medicine, question-answering agents, and education.

However, current LLMs are known to generate many kinds of unsafe system behaviors, such as providing false or inconsistent information, reporting unjustified confidence levels on rare events, or performing erroneous actions. These unsafe behaviors can lead to potentially catastrophic results in high-stakes domains, so ensuring LLM safety is a pressing question that we must address to protect against social harm.

This project focuses on enhancing the safety of LLMs by proposing quantifiable safety measures and corresponding algorithms to detect unsafe behaviors and mitigate them. Furthermore, this project will support the development of a graduate-level course on trustworthy AI, which will be offered to students from underrepresented groups to promote diversity in AI research at the University of Illinois Urbana-Champaign.

The technical aims of the project contains three key thrusts: (1) Robust-Confidence Safety (RCS), which ensures that LLMs recognize and appropriately respond to out-of-distribution scenarios and rare events; (2) Self-Consistency Safety (SCS), which enforces logical consistency in LLM outputs across similar contexts; and (3) Alignment Safety (AS), which aligns LLM responses with user objectives, particularly to avoid generating false or misleading information. The project will define these safety criteria, develop detection methods for unsafe scenarios, and create algorithms to enhance LLM safety.

The proposed methods will be tested using the open-source LLM framework LMFlow, ensuring access for creating practical applications and community availability. The project promises significant benefits, including safer AI applications, advancements in the field, and contributions to education and diversity.

This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.

All Grantees

University of Illinois At Urbana-Champaign

Advertisement
Discover thousands of grant opportunities
Advertisement
Browse Grants on GrantFunds
Interested in applying for this grant?

Complete our application form to express your interest and we'll guide you through the process.

Apply for This Grant