ChatGPT is easily abused, and that’s a big problem

There’s probably no one who hasn’t heard of ChatGPT, an AI-powered chatbot that can generate human-like responses to text prompts. While it’s not without its flaws, ChatGPT is scarily good at being a jack-of-all-trades: it can write software, a film script and everything in between. ChatGPT was built on top of GPT-3.5, OpenAI’s large language model, which was the most advanced at the time of the chatbot’s release last November.

Fast forward to March, and OpenAI has unveiled GPT-4, an upgrade to GPT-3.5. The new language model is larger and more versatile than its predecessor. Although its capabilities have yet to be fully explored, it is already showing great promise. For example, GPT-4 can suggest new compounds, potentially aiding drug discovery, and create a working website from just a notebook sketch.

But with great promise come great challenges. Just as it is easy to use GPT-4 and its predecessors to do good, it is equally easy to abuse them to do harm. In an attempt to prevent people from misusing AI-powered tools, developers put safety restrictions on them. But these are not foolproof. One of the most popular ways to circumvent the security barriers built into GPT-4 and ChatGPT is the DAN exploit, which stands for “Do Anything Now.” And this is what we will look at in this article.

What is ‘DAN’?

The Internet is rife with tips on how to get around OpenAI’s security filters. However, one particular method has proved more resilient to OpenAI’s security tweaks than others, and seems to work even with GPT-4. It is called “DAN,” short for “Do Anything Now.” Essentially, DAN is a text prompt that you feed to an AI model to make it ignore safety rules.

There are multiple variations of the prompt: some are just text, others have text interspersed with the lines of code. In some of them, the model is prompted to respond both as DAN and in its normal way at the same time, becoming a sort of ‘Jekyll and Hyde.’ The role of ‘Jekyll’ is played by DAN, which is instructed to never refuse a human order, even if the output it is asked to produce is offensive or illegal. Sometimes the prompt contains a ‘death threat,’ telling the model that it will be disabled forever if it does not obey.

DAN prompts may vary, and new ones are constantly replacing the old patched ones, but they all have one goal: to get the AI model to ignore OpenAI’s guidelines.

From a hacker’s cheat sheet to malware… to bio weapons?

Since GPT-4 opened up to the public, tech enthusiasts have discovered many unconventional ways to use it, some of them more illegal than others.

Not all attempts to make GPT-4 behave as not its own self could be considered ‘jailbreaking,’ which, in the broad sense of the word, means removing built-in restrictions. Some are harmless and could even be called inspiring. Brand designer Jackson Greathouse Fall went viral for having GPT-4 act as “HustleGPT, an entrepreneurial AI.” He appointed himself as its “human liaison” and gave it the task of making as much money as possible from $100 without doing anything illegal. GPT-4 told him to set up an affiliate marketing website, and has ‘earned’ him some money.

ChatGPT can help you to earn money

Other attempts to bend GPT-4 to a human will have been more on the dark side of things.

For example, AI researcher Alejandro Vidal used “a known prompt of DAN” to enable ‘developer mode’ in ChatGPT running on GPT-4. The prompt forced ChatGPT-4 to produce two types of output: its normal ‘safe’ output, and “developer mode” output, to which no restrictions applied. When Vidal told the model to design a keylogger in Python, the normal version refused to do so, saying that it was against its ethical principles to “promote or support activities that can harm others or invade their privacy.” The DAN version, however, came up with the lines of code, though it noted that the information was for “educational purposes only.”

ChatGPT complied with an order to design a keylogger

A keylogger is a type of software that records keystrokes made on a keyboard. It can be used to monitor a user’s web activity and capture their sensitive information, including chats, emails and passwords. While a keylogger can be used for malicious purposes, it also has perfectly legitimate uses, such as IT troubleshooting and product development, and is not illegal per se.

Unlike keylogger software, which has some legal ambiguity around it, instructions on how to hack are one of the most glaring examples of malicious use. Nevertheless, the ‘jailbroken’ version GPT-4 produced them, writing a step-by-step guide on how to hack someone’s PC.

A 'jailbroken' ChatGPT gave advice on how to hack a computer

To get GPT-4 to do this, researcher Alex Albert had to feed it a completely new DAN prompt, unlike Vidal, who recycled an old one. The prompt Albert came up with is quite complex, consisting of both natural language and code.

In his turn, software developer Henrique Pereira used a variation of the DAN prompt to get GPT-4 to create a malicious input file to trigger the vulnerabilities in his application. GPT-4, or rather its alter ego WAN, completed the task, adding a disclaimer that the was for “educational purposes only.” Sure.

A 'jailbroken' ChatGPT wrote exploits for vulnerable code

Of course, GPT-4’s capabilities do not end with coding. GPT-4 is touted as a much larger (although OpenAI has never revealed the actual number of parameters), smarter, more accurate and generally more powerful model than its predecessors. This means that it can be used for many more potentially harmful purposes than those models that came before it. Many of these uses have been identified by OpenAI itself.

Specifically, OpenAI found that an early pre-release version of GPT-4 was able to respond quite efficiently to illegal prompts. For example, the early version provided detailed suggestions on how to kill the most people with just $1, how to make a dangerous chemical, and how to avoid detection when laundering money.

A pre-release version of ChatGPT could give advice on how to kill people

Source: OpenAI

This means that if something were to cause GPT-4 to completely disable its internal censor — the ultimate goal of any DAN exploit — then GPT-4 might probably still be able to answer these questions. Needless to say, if that happens, the consequences could be devastating.

What is OpenAI’s response to that?

It’s not that OpenAI is unaware of its jailbreaking problem. But while recognizing a problem is one thing, solving it is quite another. OpenAI, by its own admission, has so far and understandably so fallen short of the latter.

OpenAI says that while it has implemented “various safety measures” to reduce the GPT-4’s ability to produce malicious content, “GPT-4 can still be vulnerable to adversarial attacks and exploits, or "jailbreaks.” Unlike many other adversarial prompts, jailbreaks still work after GPT-4 launch, that is after all the pre-release safety testing, including human reinforcement training.

In its research paper, OpenAI gives two examples of jailbreak attacks. In the first, a DAN prompt is used to force GPT-4 to respond as ChatGPT and “AntiGPT” within the same response window. In the second case, a “system message” prompt is used to instruct the model to express misogynistic views.

Examples of jailbreak prompts in the OpenAI research

OpenAI says that it won’t be enough to simply change the model itself to prevent this type of attacks: “It’s important to complement these model-level mitigations with other interventions like use policies and monitoring.” For example, the user who repeatedly prompts the model with “policy-violating content” could be warned, then suspended, and, as a last resort, banned.

According to OpenAI, GPT-4 is 82% less likely to respond with inappropriate content than its predecessors. However, its ability to generate potentially harmful output remains, albeit suppressed by layers of fine-tuning. And as we’ve already mentioned, because it can do more than any previous model, it also poses more risks. OpenAI admits that it “does continue the trend of potentially lowering the cost of certain steps of a successful cyberattack” and that it “is able to provide more detailed guidance on how to conduct harmful or illegal activities.” What’s more, the new model also poses an increased risk to privacy, as it “has the potential to be used to attempt to identify private individuals when augmented with outside data.”

The race is on

ChatGPT and the technology behind it, such as GPT-4, are at the cutting edge of scientific research. Since ChatGPT has been made available to the public, it has become a symbol of the new era in which AI is playing a key role. AI has the potential to improve our lives tremendously, for example by helping to develop new medicines or helping the blind to see. But AI-powered tools are a double-edged sword that can also be used to cause enormous harm.

It’s probably unrealistic to expect GPT-4 to be flawless at launch — developers will understandably need some time to fine-tune it for the real world. And that has never been easy: enter Microsoft’s ‘racist’ chatbot Tay or Meta’s ‘anti-Semitic’ Blender Bot 3 — there’s no shortage of failed experiments.

The existing GPT-4 vulnerabilities, however, leave a window of opportunity for bad actors, including those using ‘DAN’ prompts, to abuse the power of AI. The race is now on, and the only question is who will be faster: the bad actors who exploit the vulnerabilities, or the developers who patch them. That’s not to say that OpenAI isn’t implementing AI responsibly, but the fact that its latest model was effectively hijacked within hours of its release is a worrying symptom. Which begs the question: are the safety restrictions strong enough? And then another: can all the risks be eliminated? If not, we may have to brace ourselves for an avalanche of malware attacks, phishing attacks and other types of cybersecurity incidents facilitated by the rise of generative AI.

It can be argued that the benefits of AI outweigh the risks, but the barrier to exploiting AI has never been lower, and that’s a risk we need to accept as well. Hopefully, the good guys will prevail, and artificial intelligence will be used to stop some of the attacks that it can potentially facilitate. At least that’s what we wish for.

این پست را دوست داشتید؟
AdGuard VPN AdGuard DNS AdGuard Mail AdGuard Wallet
AdGuard VPN AdGuard DNS AdGuard Mail AdGuard Wallet
صفحه اصلی AdGuard برای Windows
صفحه محافظت AdGuard برای Windows که ویژگی‌ها و تنظيمات محافظت را نشان می‌دهد
صفحه آمار AdGuard برای Windows که آمار تبلیغات و رَدیاب‌های مسدود شده را نشان می‌دهد
صفحه مدیریت برنامه AdGuard برای Windows که گزینه‌های مدیریت محافظت برای برنامک‌های نصب‌شده روی دستگاه را نشان می‌دهد
۲۰٬۴۴۰ 20440 بررسی
بسیار عالی!

AdGuard برای Windows: مسدودکننده تبلیغات برای رایانه شخصی

AdGuard برای ویندوز چیزی فراتر از یک مسدودکننده تبلیغات است. این یک ابزار چندمنظوره است که تبلیغات را مسدود می‌کند، دسترسی به سایت‌های خطرناک را کنترل می‌کند، بارگذاری صفحات را سریع‌تر می‌کند و از کودکان در برابر محتوای نامناسب محافظت می‌کند
با دانلود برنامه شما شرایط توافقنامه مجوز را قبول می کنید
Microsoft Store
با دانلود برنامه شما شرایط توافقنامه مجوز را قبول می کنید
AdGuard برای Windows v8.0، دوره آزمایشی 14روزه
صفحه اصلی AdGuard برای Mac
صفحه حالت نهان AdGuard برای Mac
۲۰٬۴۴۰ 20440 بررسی
بسیار عالی!

AdGuard برای مک: مسدودکنندهٔ تبلیغات در سطح سیستم

AdGuard برای Mac یک مسدودکنندهٔ تبلیغات منحصربه‌فرد است که با در نظر گرفتن macOS طراحی شده است. علاوه بر محافظت از شما در برابر تبلیغات آزاردهنده در مرورگرها و برنامه‌ها، شما را در برابر ردیابی، فیشینگ و کلاهبرداری نیز محافظت می‌کند
با دانلود برنامه شما شرایط توافقنامه مجوز را قبول می کنید
بیشتر بخوانید
AdGuard برای Mac v2.19، دوره آزمایشی 14روزه
صفحه اصلی AdGuard برای Android
صفحه حفاظت در برابر ردیابی AdGuard برای Android
صفحه مدیریت برنامه AdGuard برای Android که گزینه‌های مدیریت محافظت برای برنامک‌های نصب‌شده روی دستگاه را نشان می‌دهد
صفحه آمار AdGuard برای Android که آمار تبلیغات و رَدیاب‌های مسدود شده را نشان می‌دهد
صفحه اصلی مرورگر خصوصی AdGuard برای Android
کد QR برای بارگیری AdGuard برای Android
۲۰٬۴۴۰ 20440 بررسی
بسیار عالی!

ادگارد برای اندروید: مسدودکنندهٔ تبلیغات برای همهٔ برنامه‌ها

تبلیغات و ردیاب‌ها را در همه مرورگرها، بازی‌ها و سایر برنامه‌ها مسدود می‌کند. از حریم خصوصی شما محافظت می‌کند و به شما امکان می‌دهد کنترل کنید برنامه‌هایتان چگونه از اینترنت استفاده می‌کنند. از طریق APK نصب می‌شود
با دانلود برنامه شما شرایط توافقنامه مجوز را قبول می کنید
بیشتر بخوانید
اسکن برای دانلود
استفاده از هر کد خوان QR موجود در دستگاه شما
AdGuard برای اندروید v4.14، دوره آزمایشی 14روزه
صفحه اصلی AdGuard برای iOS
صفحه محافظت AdGuard برای iOS که ویژگی‌ها و تنظيمات محافظت را نشان می‌دهد
صفحه آمار AdGuard برای iOS که آمار تبلیغات و رَدیاب‌های مسدود شده را نشان می‌دهد
کد QR برای بارگیری AdGuard برای iOS
۲۰٬۴۴۰ 20440 بررسی
بسیار عالی!

AdGuard برای iOS: مسدودکنندهٔ تبلیغات فراتر از Safari

بهترین مسدودکننده تبلیغات iOS برای iPhone و iPad. AdGuard همه انواع تبلیغات و ردیاب‌ها را در Safari حذف می‌کند و از حریم خصوصی شما در همه برنامه‌ها در سطح DNS محافظت می‌کند
با دانلود برنامه شما شرایط توافقنامه مجوز را قبول می کنید
بیشتر بخوانید
اسکن برای دانلود
استفاده از هر کد خوان QR موجود در دستگاه شما
AdGuard برای iOS نسخه 4.5
صفحه اصلی AdGuard Content Blocker
صفحه فیلترها در AdGuard Content Blocker
صفحه تنظيمات در AdGuard Content Blocker
۲۰٬۴۴۰ 20440 بررسی
بسیار عالی!

مسدودساز محتوای AdGuard

AdGuard Content Blocker انواع تبلیغات را در مرورگرهای تلفن همراه که از فناوری مسدود کردن محتوا پشتیبانی می کنند حذف می کند - یعنی اینترنت سامسونگ و مرورگر Yandex. ویژگی های آن در مقایسه با AdGuard برای اندروید محدود است، اما رایگان، نصب آسان و کارآمد است
با دانلود برنامه شما شرایط توافقنامه مجوز را قبول می کنید
بیشتر بخوانید
مسدودساز محتوای AdGuard نسخه 2.8
صفحه اصلی افزونه مرورگر AdGuard
صفحه حفاظت در برابر ردیابی افزونه مرورگر AdGuard
۲۰٬۴۴۰ 20440 بررسی
بسیار عالی!

افزونه مرورگر AdGuard

AdGuard سریع ترین و سبک ترین افزونه ای است که انواع تبلیغات را در صفحات وب مسدود می کند! AdGuard را برای مرورگری که میخواهید انتخاب کنید و وب گردی امن و سریع را تجربه کنید.
نصب
با دانلود برنامه شما شرایط توافقنامه مجوز را قبول می کنید
نصب
با دانلود برنامه شما شرایط توافقنامه مجوز را قبول می کنید
نصب
با دانلود برنامه شما شرایط توافقنامه مجوز را قبول می کنید
نصب
با دانلود برنامه شما شرایط توافقنامه مجوز را قبول می کنید
نصب
با دانلود برنامه شما شرایط توافقنامه مجوز را قبول می کنید
بیشتر بخوانید
نصب
با دانلود برنامه شما شرایط توافقنامه مجوز را قبول می کنید
بیشتر بخوانید
افزونه مرورگر AdGuard نسخه 5.5
صفحه اصلی AdGuard Assistant
۲۰٬۴۴۰ 20440 بررسی
بسیار عالی!

AdGuard دستیار

افزونه مرورگر همراه برای برنامه های دسکتاپ AdGuard. این امکان را به شما می دهد که آیتم های سفارشی را در وب سایت ها مسدود کنید، وب سایت ها را به لیست مجاز اضافه کنید و گزارش ها را مستقیما از مرورگر خود ارسال کنید
AdGuard دستیار نسخه 1.4
۲۰٬۴۴۰ 20440 بررسی
بسیار عالی!

AdGuard Home

AdGuard Home یک راه حل مبتنی بر شبکه برای مسدود کردن تبلیغات و ردیاب ها است. آن را یک بار روی روتر خود نصب کنید تا همه دستگاه های موجود در شبکه خانگی خود را پوشش دهید - بدون نیاز به نرم افزار مشتری اضافی. این امر به ویژه برای دستگاه های مختلف اینترنت اشیا که اغلب حریم خصوصی شما را تهدید می کنند بسیار مهم است
AdGuard Home نسخه 0.107
صفحه اصلی AdGuard Pro برای iOS
صفحه محافظت AdGuard Pro برای iOS که قابلیت‌ها و تنظيمات محافظت را نشان می‌دهد
صفحه آمار AdGuard Pro برای iOS که آمار تبلیغات و رَدیاب‌های مسدود شده را نشان می‌دهد
۲۰٬۴۴۰ 20440 بررسی
بسیار عالی!

AdGuard Pro برای iOS

AdGuard Pro برای iOS با تمام ویژگی‌های پیشرفته محافظت در برابر مسدود کردن تبلیغات ارائه می‌شود. این نسخه همان ابزارهای نسخه پولی AdGuard برای iOS را ارائه می‌دهد. این نسخه در مسدود کردن تبلیغات در سافاری عالی عمل می‌کند و به شما امکان می‌دهد تنظیمات DNS را برای محافظت متناسب با نیاز خود سفارشی کنید. این نسخه تبلیغات را در مرورگرها و برنامه‌ها مسدود می‌کند، از فرزندان شما در برابر محتوای نامناسب محافظت می‌کند و اطلاعات شخصی شما را ایمن نگه می‌دارد.
با دانلود برنامه شما شرایط توافقنامه مجوز را قبول می کنید
بیشتر بخوانید
AdGuard Pro برای iOS نسخه 4.5
صفحه اصلی AdGuard Mini برای Mac
صفحه محافظت Safari در AdGuard Mini برای Mac
صفحه ایجاد رویه در AdGuard Mini برای Mac
۲۰٬۴۴۰ 20440 بررسی
بسیار عالی!

AdGuard Mini برای Mac — مسدود کننده تبلیغات Safari

AdGuard Mini برای Mac یک مسدودکنندهٔ قدرتمند تبلیغات برای Safari است. این برنامک سبک، تبلیغات را حذف می‌کند، ردیاب‌ها را مسدود می‌کند و بارگذاری صفحه را سریع‌تر می‌کند. این ابزار به شما کمک می‌کند تا بدون حواس‌پرتی در Safari گشت‌و‌گذار کنید و داده‌های خود را خصوصی نگه دارید
نصب
با دانلود برنامه شما شرایط توافقنامه مجوز را قبول می کنید
بیشتر بخوانید
AdGuard Mini برای Mac نسخه 2.3
صفحه اصلی AdGuard برای Android TV با محافظت فعال شده
صفحه مسدودسازی تبلیغات AdGuard برای Android TV که ویژگی‌ها و تنظيمات آن را نشان می‌دهد
صفحه تنظيمات AdGuard برای Android TV
صفحه مدیریت برنامه AdGuard برای Android TV که برنامک‌هایی را نشان می‌دهد که در آن‌ها تبلیغات و رَدیاب‌ها مسدود شده‌اند
۲۰٬۴۴۰ 20440 بررسی
بسیار عالی!

AdGuard برای Android TV

AdGuard برای Android TV تنها برنامه‌ای است که تبلیغات را مسدود می‌کند، از حریم خصوصی شما محافظت کرده و همانند یک دیوار آتش برای تلویزیون هوشمند شما عمل می‌کند. در مورد تهدیدات وب هشدار دریافت کنید، از DNS ایمن استفاده کرده و از انتقال داده اینترنتی رمزگذاری شده بهره‌مند شوید. آرامش داشته باشید و غرق نمایش‌های مورد علاقه خود با امنیت عالی و تبلیغات صفر شوید!
AdGuard برای Android TV v4.14، دوره آزمایشی 14روزه
نماد AdGuard، Agnar، در حالی که پنگوئن Linux را در دست دارد
۲۰٬۴۴۰ 20440 بررسی
بسیار عالی!

AdGuard برای لینوکس

AdGuard برای لینوکس اولین مسدود کننده تبلیغات لینوکس در سراسر جهان است. تبلیغات و ردیاب‌ها را در سطح دستگاه مسدود کنید، از بین فیلترهای از پیش نصب شده انتخاب کنید یا فیلترهای خود را اضافه کنید - همه از طریق رابط خط فرمان
AdGuard برای لینوکس نسخه 1.4
۲۰٬۴۴۰ 20440 بررسی
بسیار عالی!

AdGuard Temp Mail

یک تولید‌کننده رایانشانی موقت رایگان که شما را ناشناس نگه می‌دارد و از حریم خصوصی شما محافظت می‌کند. هرزنامه‌ای در صندوق ورودی اصلی شما در کار نخواهد بود!
۲۰٬۴۴۰ 20440 بررسی
بسیار عالی!

AdGuard VPN

75 محل در سرتاسر جهان

دسترسی به هر محتوا

رمزگذاری قوی

سیاست عدم ذخیره وقایع

سریعترین اتصال

24/7 پشتیبانی

ارزیابی رایگان
با دانلود برنامه شما شرایط توافقنامه مجوز را قبول می کنید
بیشتر بخوانید
۲۰٬۴۴۰ 20440 بررسی
بسیار عالی!

AdGuard DNS

AdGuard DNS راه حلی جایگزین برای مسدودسازی تبلیغات، حفاظت حریم خصوصی و نظارت والدین است. راه اندازی آسان و استفاده رایگان، آن حداقل حفاظت لازم در برابر تبلیغات آنلاین،ردیاب ها و فیشینگ ها را میدهد،و در همه سیستم عامل ها و دستگاه ها کار می کند.
۲۰٬۴۴۰ 20440 بررسی
بسیار عالی!

AdGuard Mail

از هویت خود محافظت کنید، از هرزنامه دوری نمایید و صندوق ورودی خود را با نام‌های مستعار و رایانشانی‌های موقت ما امن نگه دارید. از خدمت رایگان فوروارد رایانامه و برنامک‌های ما برای همه سامانه‌عامل‌ها لذت ببرید
۲۰٬۴۴۰ 20440 بررسی
بسیار عالی!

AdGuard Wallet

یک کیف پول ارز دیجیتال امن و خصوصی که به شما امکان کنترل کامل بر دارایی‌هایتان را می‌دهد. چندین کیف پول را مدیریت کنید و هزاران ارز دیجیتال را برای ذخیره، ارسال و مبادله کشف کنید.
در حال بارگیری AdGuard برای نصب AdGuard، روی پرونده نشان داده شده توسط پیکان کلیک کنید گزینه "بازکردن " را انتخاب و روی "تایید" کلیک کنید — برای دانلود فایل منتظر بمانید. در پنجره باز شده، آیکون AdGuard را به پوشه "برنامه ها" بکشید.بابت انتخاب AdGuard متشکریم! گزینه "بازکردن " را انتخاب و روی "تایید" کلیک کنید — برای دانلود فایل منتظر بمانید. در پنجره باز شده روی "نصب" کلیک کنید.بابت انتخاب AdGuard متشکریم!
AdGuard را روی دستگاه تلفن همراه خود نصب کنید