<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Jailbreak on CuraSec</title><link>https://curasec.metacog.co.kr/tags/jailbreak/</link><description>Recent content in Jailbreak on CuraSec</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Mon, 31 Aug 2026 19:07:02 +0000</lastBuildDate><atom:link href="https://curasec.metacog.co.kr/tags/jailbreak/index.xml" rel="self" type="application/rss+xml"/><item><title>Mechanistic Interpretability Study Maps LLM Jailbreak Circuits</title><link>https://curasec.metacog.co.kr/insights/2026-08-31-circuit-discovery-helps-detect-llm-jailbreaking-a-mechanisti/</link><pubDate>Mon, 31 Aug 2026 19:07:02 +0000</pubDate><guid>https://curasec.metacog.co.kr/insights/2026-08-31-circuit-discovery-helps-detect-llm-jailbreaking-a-mechanisti/</guid><description>&lt;ul>
&lt;li>&lt;strong>Engineer — Learn:&lt;/strong> Research identifies internal attention heads and MLP pathways responsible for safety bypass in LLaMA-2-7B — useful context when evaluating LLM safeguard architectures, but no operational change needed today and findings are on one specific model.&lt;/li>
&lt;li>&lt;strong>SOC/IR — Skip&lt;/strong>&lt;/li>
&lt;li>&lt;strong>Leader — Learn:&lt;/strong> Findings suggest current LLM safety alignment has exploitable structural weaknesses; relevant context when assessing risk posture of internally deployed LLM products, but no immediate action required.&lt;/li>
&lt;/ul></description></item><item><title>Transferable CoT Jailbreaks Bypass LLM Output Safeguards at Scale</title><link>https://curasec.metacog.co.kr/insights/2026-07-20-hidden-in-thought-transferable-chain-of-thought-artifacts-in/</link><pubDate>Mon, 20 Jul 2026 14:31:24 +0000</pubDate><guid>https://curasec.metacog.co.kr/insights/2026-07-20-hidden-in-thought-transferable-chain-of-thought-artifacts-in/</guid><description>&lt;ul>
&lt;li>&lt;strong>Engineer — Learn:&lt;/strong> Research shows output-only filters like Llama-Guard 3 are insufficient against reasoning-layer attacks; teams building AI applications should evaluate reasoning context, not just final outputs, when designing safety architectures.&lt;/li>
&lt;li>&lt;strong>SOC/IR — Skip&lt;/strong>&lt;/li>
&lt;li>&lt;strong>Leader — Learn:&lt;/strong> Finding that reasoning-capable models are 2x+ more vulnerable and standard output safeguards regularly fail has implications for enterprise AI risk posture; useful context for AI usage policies and vendor safety attestation reviews.&lt;/li>
&lt;/ul></description></item></channel></rss>