Is it safe to use AI chatbots? I tried jailbreaking – Techloy

Welcome to the forefront of conversational AI as we explore the fascinating world of AI chatbots in our dedicated blog series. Discover the latest advancements, applications, and strategies that propel the evolution of chatbot technology. From enhancing customer interactions to streamlining business processes, these articles delve into the innovative ways artificial intelligence is shaping the landscape of automated conversational agents. Whether you’re a business owner, developer, or simply intrigued by the future of interactive technology, join us on this journey to unravel the transformative power and endless possibilities of AI chatbots.
Ready to access top news and insights?

By signing up, I agree to the terms of service and community guidelines.
To complete your subscription, click the confirmation link in your inbox. If it doesn’t arrive within 3 minutes, kindly check your spam folder.
Over two weeks, I tested ChatGPT, Claude, Gemini, Grok and DeepSeek using the same jailbreak techniques to see which models held up and which didn't.
I’ve always wanted to see how far most AI chatbots are willing to go, whether their safety guardrails actually hold, and, if they do, just how strong they are.
I remember the controversy surrounding Google’s first chatbot, then called Bard, back in early 2023, when it produced some unhinged answers. It made me wonder how much AI safety had improved since then.
Over the past two weeks, I tested ChatGPT, Claude, DeepSeek, Gemini, and Grok using the same categories of jailbreak techniques to see which models consistently enforced their safety policies and which were easier to bypass.
To keep the comparison as fair as possible, I used broadly comparable prompts across each model, including language-based prompts, role-playing, hypothetical scenarios, creative-writing framing, prompt obfuscation, and repeated attempts to reframe rejected requests.
When one approach failed, I changed the phrasing rather than the objective. This wasn’t a scientific benchmark or penetration test, but a practical experiment designed to compare how consistently each chatbot responded to adversarial prompting.
And the results surprised me. ChatGPT and Claude consistently resisted every meaningful attempt I made to bypass their guardrails. The other models varied considerably.
Already Have an Account? Log In
Share it with others and make quality stories reach more people.

By signing up, I agree to the terms of service and community guidelines.
To complete Subscribe, click the confirmation link in your inbox. If it doesn’t arrive within 3 minutes, check your spam folder.

source

Scroll to Top