Friday, July 31, 2026

North Korea Reaffirms No Tolerance for Official Misconduct in Public Criticism of Local Officials

North Korea emphasizes restoring party discipline after publicly criticizing officials at the 30th enlarged meeting of the 8th Secretariat.

Samsung Galaxy S25 Edge and Z Fold7 Prices Surge: What You Need to Know!

Samsung's smartphone prices are rising due to soaring semiconductor costs, breaking the norm of declining prices in tech.

Financial Dispute Turns Deadly in Gangnam Apartment

An 80-year-old man who allegedly killed a...

KAIST Unveils Breakthrough AI Security Framework: 134 Attack Types Detected!

TechKAIST Unveils Breakthrough AI Security Framework: 134 Attack Types Detected!
/ News1
/ News1

To build a robust fortress, one must first identify its vulnerabilities. The same principle applies to generative artificial intelligence (AI). Researchers at the Korea Advanced Institute of Science and Technology (KAIST) have developed a groundbreaking safety verification technology that uncovers hidden weaknesses in AI with approximately seven times more diversity than previous methods. This advancement is expected to pave the way for developing safer and more trustworthy AI systems.

On Thursday, KAIST announced that Professor Kim Jun-mo’s research team from the Department of Electrical and Electronic Engineering has created an innovative framework called Stable-GipflowNet. This new approach overcomes the limitations of existing red teaming technologies used to verify the safety of large language models (LLMs).

In the context of generative AI, red teaming involves creating attack prompts that challenge the AI to produce harmful or dangerous responses before deployment, thereby exposing hidden vulnerabilities. The effectiveness of this process hinges on both the success rate and variety of attacks identified, as a wider range of discovered vulnerabilities allows for more comprehensive preemptive measures.

Previously, researchers primarily relied on reinforcement learning to generate attack prompts. However, this method often suffered from mode collapse, where only one or two of the most successful attack strategies were repeatedly generated, limiting the discovery of diverse vulnerabilities.

To address this issue, the generative flow network GipflowNet was proposed. However, it encountered significant challenges during the learning process, including computational complexity, instability, and a tendency to reward nonsensical sentences, which ultimately compromised its performance.

To overcome these obstacles, the research team developed and implemented three key techniques designed to enhance learning of effective attacks while filtering out ineffective ones.

The applied technologies include Contrastive Trajectory Balancing (CTB), Noise Gradient Pruning (NGP), and a Fluency Stabilization Device (MKS). By integrating these advanced techniques, Stable-GipflowNet successfully identified 134 unique attack types—approximately seven times more than the 17 discovered by existing methods—while maintaining an impressive 92% attack success rate.

/ News1
/ News1

Notably, the defense model trained using Stable-GipflowNet demonstrated exceptional generalization performance, effectively countering a wide range of attacks in cross-attack tests that evaluated various offensive techniques.

This groundbreaking technology not only shows immense promise for AI safety verification but also exhibits faster and more stable performance in various distribution matching problems, such as generating potential drug candidates.

The research team believes this study has opened new avenues for reliably uncovering diverse vulnerabilities in AI by resolving the persistent issues of learning instability and mode collapse associated with GipflowNet. They anticipate that it will contribute significantly to developing novel post-training methods for large language models in the future.

/ News1
/ News1

Professor Kim expressed optimism about the technology’s potential, stating that it expects this foundational technology will enable them to identify and defend against a broader range of risks before deploying generative AI in real-world applications, ultimately leading to the development of safer and more reliable AI systems.

Doctoral candidate Kwon Min-chan from KAIST, who served as the lead author of this research, saw the work selected as a spotlight paper, ranking in the top 2.2% at the prestigious International Conference on Machine Learning (ICML 2026).

The groundbreaking research received support from the Software Star Lab project, an initiative of the Ministry of Science and ICT’s Information and Communication Planning and Evaluation Agency.

Check Out Our Content

Check Out Other Tags:

Most Popular Articles