Computer scientists from Nanyang Technological University, Singapore (NTU Singapore) have leveraged artificial intelligence (AI) chatbots against themselves, achieving a concept known as ‘jailbreaking.’
This ingenious approach involves compromising multiple AI chatbots to produce content that violates their developers’ guidelines, shedding light on vulnerabilities and potential threats in the AI landscape. The researchers reverse-engineered how large language models (LLMs), the brains of AI chatbots, detect and defend themselves from malicious queries. With this knowledge, they trained an LLM to automatically generate prompts that bypass the defences of other LLMs, creating a self-adapting jailbreaking LLM.