<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Vision-Language-Models on CuraSec</title><link>https://curasec.metacog.co.kr/tags/vision-language-models/</link><description>Recent content in Vision-Language-Models on CuraSec</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Mon, 31 Aug 2026 19:07:02 +0000</lastBuildDate><atom:link href="https://curasec.metacog.co.kr/tags/vision-language-models/index.xml" rel="self" type="application/rss+xml"/><item><title>Meta-Adaptive Jailbreaking Achieves 80%+ ASR Against Frontier VLMs</title><link>https://curasec.metacog.co.kr/insights/2026-08-31-fully-unleashing-the-multimodal-attacker-meta-adaptive-jailb/</link><pubDate>Mon, 31 Aug 2026 19:07:02 +0000</pubDate><guid>https://curasec.metacog.co.kr/insights/2026-08-31-fully-unleashing-the-multimodal-attacker-meta-adaptive-jailb/</guid><description>&lt;ul>
&lt;li>&lt;strong>Engineer — Learn:&lt;/strong> If your product integrates GPT-4o, Gemini, or similar multimodal models, this research shows existing content-safety wrappers are brittle against adaptive attackers; no patch exists yet, but it motivates evaluating your VLM endpoints against adaptive prompt-injection test suites.&lt;/li>
&lt;li>&lt;strong>SOC/IR — Skip&lt;/strong>&lt;/li>
&lt;li>&lt;strong>Leader — Learn:&lt;/strong> Research demonstrating high-success jailbreaks against GPT-4o and Gemini is useful framing for board-level AI risk discussions and for questioning AI vendor safety assurance claims when procuring or expanding VLM-based tooling.&lt;/li>
&lt;/ul></description></item><item><title>Adversarial Attacks Can Silently Manipulate VLM Confidence Scores</title><link>https://curasec.metacog.co.kr/insights/2026-08-10-model-confidence-under-answer-preserving-attacks-an-informat/</link><pubDate>Mon, 10 Aug 2026 13:39:41 +0000</pubDate><guid>https://curasec.metacog.co.kr/insights/2026-08-10-model-confidence-under-answer-preserving-attacks-an-informat/</guid><description>&lt;ul>
&lt;li>&lt;strong>Engineer — Learn:&lt;/strong> White-box attacks can degrade confidence readouts in vision-language models to near-random while leaving the generated answer unchanged, undermining confidence-gated pipelines; teams deploying VLMs with confidence thresholds for access control or oversight should treat confidence as an untrusted signal in adversarial contexts.&lt;/li>
&lt;li>&lt;strong>SOC/IR — Skip&lt;/strong>&lt;/li>
&lt;li>&lt;strong>Leader — Learn:&lt;/strong> Academic research showing that AI confidence gating — a common oversight mechanism in deployed vision-language products — can be silently subverted; worth tracking as AI governance frameworks and internal AI-use policies mature, but no immediate action warranted.&lt;/li>
&lt;/ul></description></item></channel></rss>