Home Tech Revolutionary AI Breakthrough: How KAIST’s New Tech Reduces Sensory Confusion in Multimodal...

Revolutionary AI Breakthrough: How KAIST’s New Tech Reduces Sensory Confusion in Multimodal AI

0
/ News1
/ News1

A breakthrough technology has been developed to end the era where artificial intelligence (AI) claims to hear the inaudible and see the invisible.

On Friday, researchers at the Korea Advanced Institute of Science and Technology (KAIST) announced the development of two key technologies aimed at reducing sensory confusion in multimodal large language models (MLLMs).

This innovation significantly enhances AI reliability without the need for extensive retraining by accurately interpreting data from cameras and various sensors, while also minimizing hallucinations that occur when visual and auditory information becomes intertwined.

The first technology optimizes a method called Sensor Understanding AI (DNA), which enables AI to accurately comprehend the physical data captured by various sensors, including thermal images, depth maps, and X-rays.

While humans easily recognize bright areas in thermal images as hot spots, existing AI sometimes misinterprets these as mere light, similar to regular photographs. The research team introduced a novel learning technique that allows AI to understand physical information measured by sensors, such as heat and distance, enabling more accurate object identification even in low-visibility conditions like darkness or smoke.

/ News1
/ News1

The second technology, dubbed Sensory Confusion Prevention (MAD), reduces hallucinations caused by the interplay of visual and auditory information.

For example, when closed-circuit television (CCTV) footage shows a car passing by without recorded sound, existing multimodal AI might erroneously report hearing engine noise simply because it sees the car. The team’s solution teaches AI to first determine whether to base its judgment on visual or auditory cues, then focus on the most relevant sensory data. This approach effectively reduces sensory confusion and hallucinations without requiring complete AI retraining.

Instead of rebuilding AI from scratch using vast computing resources, the team designed DNA technology to enhance performance with minimal data, while MAD can be applied directly to existing AI systems like a software update. This method dramatically cuts time and costs while boosting the reliability of existing multimodal AI, making it ideal for swift deployment in industrial applications.

The researchers noted that allowing AI to choose which sensory input to trust significantly reduces hallucinations. They also found that teaching AI why an answer is incorrect is far more effective in reducing sensory errors than simply providing correct answers.

This technology has potential applications in autonomous vehicles that need to accurately detect pedestrians and other vehicles in low-visibility conditions, rescue robots searching for survivors in smoke-filled environments, and drones utilizing thermal cameras. It’s also expected to enhance the accuracy and reliability of AI systems analyzing X-ray images at airport security checkpoints or integrating various medical imaging modalities such as computed tomography (CT), magnetic resonance imaging (MRI), and X-rays in hospitals.

/ News1
/ News1

Professor Noh emphasized that for multimodal AI to be effective in real-world scenarios, it must accurately interpret data from various sensors without confusing multiple sensory inputs. This research establishes a foundation for reliable multimodal AI that can be trusted in everyday life and industrial settings by reducing sensory biases and hallucinations without extensive retraining.

The research was led by doctoral candidate Jeong Sang-yoon from KAIST’s Department of Electrical and Electronic Engineering. Dr. Yu Young-jun contributed as a co-first author in the DNA research.

The MAD study was presented at the International Conference on Computer Vision and Pattern Recognition (CVPR) in June, while the DNA research was published in IEEE Transactions on Image Processing, a leading journal in the field of image processing.

NO COMMENTS

LEAVE A REPLY

Please enter your comment!
Please enter your name here

Exit mobile version