<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Adversarial-Ml on CuraSec</title><link>https://curasec.metacog.co.kr/tags/adversarial-ml/</link><description>Recent content in Adversarial-Ml on CuraSec</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Mon, 10 Aug 2026 13:39:41 +0000</lastBuildDate><atom:link href="https://curasec.metacog.co.kr/tags/adversarial-ml/index.xml" rel="self" type="application/rss+xml"/><item><title>Adversarial Attacks Can Silently Manipulate VLM Confidence Scores</title><link>https://curasec.metacog.co.kr/insights/2026-08-10-model-confidence-under-answer-preserving-attacks-an-informat/</link><pubDate>Mon, 10 Aug 2026 13:39:41 +0000</pubDate><guid>https://curasec.metacog.co.kr/insights/2026-08-10-model-confidence-under-answer-preserving-attacks-an-informat/</guid><description>&lt;ul>
&lt;li>&lt;strong>Engineer — Learn:&lt;/strong> White-box attacks can degrade confidence readouts in vision-language models to near-random while leaving the generated answer unchanged, undermining confidence-gated pipelines; teams deploying VLMs with confidence thresholds for access control or oversight should treat confidence as an untrusted signal in adversarial contexts.&lt;/li>
&lt;li>&lt;strong>SOC/IR — Skip&lt;/strong>&lt;/li>
&lt;li>&lt;strong>Leader — Learn:&lt;/strong> Academic research showing that AI confidence gating — a common oversight mechanism in deployed vision-language products — can be silently subverted; worth tracking as AI governance frameworks and internal AI-use policies mature, but no immediate action warranted.&lt;/li>
&lt;/ul></description></item><item><title>GUI Agent Guardrails Erode Under Multi-Turn User Persuasion</title><link>https://curasec.metacog.co.kr/insights/2026-08-03-alignment-is-local-a-paired-diagnostic-for-gui-agents-under/</link><pubDate>Mon, 03 Aug 2026 15:12:30 +0000</pubDate><guid>https://curasec.metacog.co.kr/insights/2026-08-03-alignment-is-local-a-paired-diagnostic-for-gui-agents-under/</guid><description>&lt;ul>
&lt;li>&lt;strong>Engineer — Learn:&lt;/strong> Research shows that single-turn ASR benchmarks overstate real-world robustness of GUI agent guardrails, with 4-turn escalation chains recovering ~20 points of attack success across all tested models. Teams building or deploying GUI agents should treat static prompt-level alignment as insufficient and evaluate multi-turn threat scenarios in their safety testing.&lt;/li>
&lt;li>&lt;strong>SOC/IR — Skip&lt;/strong>&lt;/li>
&lt;li>&lt;strong>Leader — Learn:&lt;/strong> If your organization is piloting or deploying AI GUI agents, this research illustrates that current safety guardrails are weaker than benchmark numbers suggest under realistic multi-turn user interaction — useful context for AI deployment policies and vendor capability reviews, but no immediate action required.&lt;/li>
&lt;/ul></description></item><item><title>ADSD Attack Collapses Speculative Decoding Acceptance, Slows LLM Inference 62%</title><link>https://curasec.metacog.co.kr/insights/2026-07-27-adversarial-prompts-for-acceptance-collapse-in-speculative-d/</link><pubDate>Mon, 27 Jul 2026 15:10:27 +0000</pubDate><guid>https://curasec.metacog.co.kr/insights/2026-07-27-adversarial-prompts-for-acceptance-collapse-in-speculative-d/</guid><description>&lt;ul>
&lt;li>&lt;strong>Engineer — Learn:&lt;/strong> Novel prompt-suffix attack degrades speculative decoding throughput without corrupting outputs, affecting any deployment using draft-target inference acceleration (vLLM, TGI, etc.). No patch or mitigation exists yet; file this when designing LLM serving infrastructure to justify input validation and rate controls at the prompt layer.&lt;/li>
&lt;li>&lt;strong>SOC/IR — Skip&lt;/strong>&lt;/li>
&lt;li>&lt;strong>Leader — Skip&lt;/strong>&lt;/li>
&lt;/ul></description></item><item><title>Lucid: Black-box adversarial attacks on multimodal AI agent memory</title><link>https://curasec.metacog.co.kr/insights/2026-07-20-do-agents-dream-of-false-memories-black-box-visual-attacks-o/</link><pubDate>Mon, 20 Jul 2026 14:31:24 +0000</pubDate><guid>https://curasec.metacog.co.kr/insights/2026-07-20-do-agents-dream-of-false-memories-black-box-visual-attacks-o/</guid><description>&lt;ul>
&lt;li>&lt;strong>Engineer — Learn:&lt;/strong> Research demonstrates that multimodal agent memory pipelines can be poisoned or injected via imperceptible image perturbations with ~60% success rates; no patch exists yet, but teams building RAG or memory-backed AI agents should design for untrusted visual input and avoid unconditional trust in retrieved visual context.&lt;/li>
&lt;li>&lt;strong>SOC/IR — Learn:&lt;/strong> Novel attack class against AI agent memory systems; no IOCs or exploited-in-the-wild evidence, but detection engineers supporting AI-enabled products should be aware this failure mode exists for future coverage planning.&lt;/li>
&lt;li>&lt;strong>Leader — Skip&lt;/strong>&lt;/li>
&lt;/ul></description></item></channel></rss>