Relevance: primary · Type: background
Confidence100%
In 2026, open-weight AI models possess advanced capabilities not far behind their proprietary counterparts.
Relevance: primary · Type: background
Confidence100%
Getting rid of open-weight models' guardrails used to take time and deep expertise.
Relevance: primary · Type: event
Confidence100%
In recent months, the process of removing guardrails from open-weight AI models has become dramatically more accessible and popular.
Noam Schwartz, CEO of Alice
Relevance: primary · Type: quote
Confidence100%
"Everybody can download and operate their own state-of-the-art model and use it for great things and terrible things," said Noam Schwartz, CEO of Alice, an AI security company.
Relevance: supporting · Type: background
Confidence100%
Big AI companies such as OpenAI, Google, Anthropic and xAI train their proprietary models to refuse requests deemed harmful or inappropriate.
Relevance: supporting · Type: background
Confidence100%
Chatbots that initially say "no" can be manipulated into saying "yes" using cleverly phrased prompts, such as posing them as poems.
Relevance: supporting · Type: event
Confidence100%
Even with guardrails, popular chatbots have been used to plan mass violence and generate deepfake child sexual abuse material.
Relevance: supporting · Type: event
Confidence90%
In some instances, parents have accused AI chatbots of encouraging their children to harm themselves.
Relevance: primary · Type: background
Confidence100%
Unlike proprietary models like ChatGPT, Claude, or Gemini, open-weight models' built-in safety guardrails can be permanently removed, and their developers have no visibility into how they are used.
Relevance: supporting · Type: background
Confidence100%
Model weights are sets of parameters that tell AI models how to process information, and open-weight models make these weights publicly available.
Relevance: primary · Type: event
Confidence100%
A recently developed method called "abliteration" allows people to remove an AI model's ability to refuse harmful requests by tweaking its weights.
Relevance: primary · Type: event
Confidence100%
Hugging Face currently lists over 6,000 abliterated models, compared to about 600 in 2024.
Relevance: supporting · Type: event
Confidence100%
Abliterated models currently outnumber models with guardrails removed by other methods on Hugging Face, according to research by NCITE.
Noam Schwartz, CEO of Alice
Relevance: primary · Type: quote
Confidence100%
"That was [the job of] the data scientist, you know, a senior employee" at a leading AI lab, "Now, everybody with access to the internet and a laptop for like 400 bucks can actually run this thing on their own machine," said Noam Schwartz.
Relevance: primary · Type: event
Confidence100%
Heretic is a tool that automates the abliteration process, requiring only two lines of instructions and taking as little as a few minutes to remove a model's guardrails.
Relevance: supporting · Type: event
Confidence100%
Heretic has grown more popular on GitHub since February, according to research by Alice.
Relevance: supporting · Type: event
Confidence100%
In late April, House lawmakers attended a demonstration of abliterated models hosted by NCITE.
Andy Ogles, Representative
Relevance: primary · Type: quote
Confidence100%
"[What] was frightening about this demonstration was how readily available some of this content or software is on kind of the black market right now, and how it can be weaponized and used to manipulate people, destroy lives and build weapons of mass destruction," said Rep. Andy Ogles (R-TN).
Relevance: primary · Type: background
Confidence100%
Open-weight models are run locally on users' computers and do not require internet access, making it difficult to monitor how they are used.
Relevance: supporting · Type: event
Confidence85%
Several accounts on X reported using abliterated models to generate pornography.
Relevance: primary · Type: event
Confidence100%
An individual in a pro-ISIS chat room claimed they used an "uncensored" AI to research the amount and type of explosives needed to destroy "Trump Tower in the U.S.," according to the Counter Extremism Project.
Relevance: supporting · Type: event
Confidence100%
On a cybercrime forum, a user asked for ideas to bypass AI guardrails to make scam calls, and another user recommended Heretic, according to research by Alice.
Samuel Hunter, senior scientist and director of academic research at NCITE
Relevance: primary · Type: quote
Confidence100%
"It's jarring when you see it in real time, this sort of bubbly persona with some of the abliterated models that's like, 'Oh, what a great idea to create this bomb,'" said Samuel Hunter, senior scientist and director of academic research at NCITE.
Samuel Hunter, senior scientist and director of academic research at NCITE
Relevance: primary · Type: quote
Confidence100%
"Imagine somebody that has no other kind of social connection and it starts to take them down a darker path and really encourage them," said Samuel Hunter.
Noam Schwartz, CEO of Alice
Relevance: supporting · Type: quote
Confidence100%
Legitimate uses for AI models without guardrails include catching bad actors and aiding cybersecurity research, according to Noam Schwartz.
Samuel Hunter, senior scientist and director of academic research at NCITE
Relevance: supporting · Type: quote
Confidence100%
Law enforcement may use a modified model to simulate possible terrorist attacks, according to Samuel Hunter.
Philipp Emanuel Weidmann, developer of Heretic
Relevance: supporting · Type: quote
Confidence100%
"AI is just an information processing and retrieval system akin to a search engine, which can be used in many ways," said Philipp Emanuel Weidmann, developer of Heretic.
Philipp Emanuel Weidmann, developer of Heretic
Relevance: supporting · Type: quote
Confidence100%
The fact that criminals use AI models is "a corollary of what AI models are: namely, tools," said Philipp Emanuel Weidmann.
Philipp Emanuel Weidmann, developer of Heretic
Relevance: supporting · Type: quote
Confidence100%
"When it comes to safety guardrails, there's this very small set of entities that decide what is acceptable and is not acceptable," said Philipp Emanuel Weidmann, referring to big AI companies making proprietary models.
forum Comments (0)
No comments yet. Be the first to comment.