Ablating a language model refers to altering parts of the system that prevent it from responding to requests it deems inappropriate, dangerous, or harmful. These protections are built into the model during training, which involves feeding it massive amounts of data—often including information from the entire internet. This process helps the model learn and store knowledge in its complex network of parameters, which are essentially the "weights" assigned to different pieces of information. Once the initial training is complete, the model undergoes fine-tuning and alignment, where it learns to distinguish between acceptable and unacceptable behaviors based on examples. This helps it understand the context and intent behind user prompts and responses. After training, the model continues to be adjusted by modifying the weights of its neurons, allowing it to adapt to new tasks or environments. Researchers in 2024 described how the model evaluates the risk of a question or response by passing a vector through all its layers. If the risk exceeds a certain threshold, the model will refuse to answer. Some systems use an external classifier, like Shieldstral from Mistral, to detect risks in the user's prompt before it even reaches the core of the language model. If the prompt is deemed safe, it proceeds to the core, where internal protections further review the response before it is sent back to the user. These protections are designed to prevent the model from assisting in harmful activities, such as creating weapons, chemical substances, engaging in self-harm, committing fraud, launching cyberattacks, or generating malware. However, there are methods to bypass these safeguards. The first involves "jailbreaking," which means tweaking the prompt to avoid triggering the model's defenses without altering the model itself. The second method, ablation, involves directly modifying the model to remove its inhibitions. The third and most intensive method requires retraining the model on data that doesn’t include any refusal signals, effectively removing the protections entirely.