Agentic AI safety & security evaluation
Evaluating the safety of tool-using language-model agents. Ongoing work is presented at the topic and status level only.
My work asks how technical risk can be turned into structured evidence, how that evidence should be evaluated, and how findings can become controls that people can review and use.
Ongoing independent research on safety evaluation for tool-using language-model agents.
Only the title, role, period, location, and general topic are public. Methods, benchmarks, models, prompts, traces, code, data, findings, and manuscripts are not disclosed.
Each line asks how evidence can make AI-system safety more legible, repeatable, and operational.
Evaluating the safety of tool-using language-model agents. Ongoing work is presented at the topic and status level only.
Turning public risk principles and technical failure models into structured tests, evidence, deployment decisions, and mitigation checks.
Building evaluation workflows for toxic, xenophobic, misleading, and biased content across multilingual settings and document formats.
This is a public synthesis of methods used across completed work. It is not a description of the private design of the ongoing agent-safety study.
Translate a broad safety concern into explicit system boundaries, failure categories, evaluation questions, and decision criteria.
Threat models · Risk taxonomies · Standards mappingConnect test cases, model or system behavior, and multidimensional measures so findings can be reproduced and compared within a defined scope.
Adversarial testing · Experimental design · Evaluation metricsBuild data, classification, and review workflows that make model outputs legible to researchers, operators, and decision-makers.
NLP pipelines · Human review · Failure classificationCarry evidence into safety cases, mitigation verification, product requirements, and workflows that can be used beyond a single analysis.
Safety cases · Control checks · Workflow designWenlan Liang, Ruizhen Pan, Yilin Cao, Wentong Wang, Shixuan Lin, and Yujia Zhao
IEIS 2022, Lecture Notes in Operations Research, pp. 152–161. Springer Nature Singapore.
DOI 10.1007/978-981-99-3618-2_15This site links to the DOI record and does not reproduce or host the paper.