AI Safety Testing Framework for LLM and Agentic Systems
A four-layer evaluation framework that translates public AI risk guidance into testing requirements, evidence structures, deployment decisions, and mitigation checks.
Four public case studies trace my work across AI safety assurance, multilingual NLP, quantitative modelling, and production operations. Each page records the problem, role, system, methods, and public scope.
See the research contextThe portfolio spans research frameworks, applied AI, statistical models, and deployed software—showing how analysis moves into usable systems.
A four-layer evaluation framework that translates public AI risk guidance into testing requirements, evidence structures, deployment decisions, and mitigation checks.
A multilingual NLP system combining GPT-4 and a fine-tuned BERT model to support structured review of xenophobic language, toxicity, misinformation, and biased framing in media content.
A quantitative research project using ten years of Chinese bond-market data and six time-series models to analyze and forecast interest-rate behavior.
End-to-end product, workflow, data, backend, and deployment ownership for a B2B platform spanning a WeChat Mini Program, back-office system, and company website.
These are case-study records, not an article library. Research papers, private datasets, working files, and proprietary system details remain off the site.
Independent Researcher and Project Lead
Four evaluation layers · Seven trustworthiness dimensions · Three case-study systems · Five practical artifact types
Independent Researcher · UNICC-Sponsored NYU Applied Project
Three integrated analysis modules · PDF, Word, and plain text · GPT-4 + fine-tuned BERT · Streamlit web application
First Author
Ten years · Six time-series models · Chinese bond market · First author
Product and Operations Manager
Mini Program, back office, website · Five connected business domains · 34-page Mini Program · Requirements through adoption