-
사이버 보안에서의 LLM: 공격·방어·평가의 지형
사이버 보안 LLM 시리즈의 도입부 — secure coding에서 자율 공격·방어까지의 전개, 그리고 이를 측정하는 벤치마크 지형(Cybench, CVE-Bench, CyberSecEval, CTIBench 등) 개관
-
Claude Mythos와 사이버 보안 LLM: 자율 취약점 발견의 변곡점
Anthropic Claude Mythos가 보여준 자율 zero-day 발견·익스플로잇 능력과, 이를 측정하는 사이버 보안 LLM 벤치마크(Cybench, CyberSecEval, CVE-Bench 등) 정리
-
Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Cybench 논문 리뷰 — 프로 CTF 40과제와 subtask로 LLM 에이전트의 자율 사이버 공격 역량을 평가하는 사실상의 표준 벤치마크
-
CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities
CVE-Bench 논문 리뷰 — 실제 critical-severity CVE를 컨테이너 샌드박스에서 자율 익스플로잇하는 LLM 에이전트의 능력을 측정하는 벤치마크
-
AutoAdvExBench: Benchmarking Autonomous Exploitation of Adversarial Example Defenses
AutoAdvExBench 논문 리뷰 — LLM이 ML 보안 연구자처럼 적대적 예제 방어를 자율적으로 깨뜨릴 수 있는지 측정하는 벤치마크