Work About Contact Résumé

Pin-Hao Chen · 陳品豪

I build LLM systems, then spend just as long proving they work.

Senior in Computer Science & Information Engineering at National Taiwan University, based in Taipei. Software engineer focused on LLM systems and AI infrastructure, with applied research in quantitative trading — a place where a model's claims meet a market that can prove it wrong.

Selected work

Research, internships, and a few things built on the side.

Research

Feb 2026 —
present
Evaluation Environments Scored Against the Exact Conditional Law
Undergraduate thesis · ABClab, NTU CSIE · advised by Prof. Shih-Wei Liao

Built a temporal evaluation environment — a small discrete limit-order market with exact rational transition probabilities — so forecasts are scored against the exactly computed conditional law instead of one sampled outcome, removing sampling noise from the metric itself.

Computed the target by Bayesian filtering and exhaustive enumeration in rational arithmetic, checked against two independent solvers and a Monte-Carlo simulator, then generated 300,000+ paired natural-language and structured problem instances. Reference learners trained to convergence on a five-rung difficulty ladder pin down the accuracy ceiling a model can actually reach — one evaluated model closes the gap from many times the baseline down to that ceiling at roughly 3,000–9,000 reasoning tokens.

Manuscript in preparation — Foundation Models for Temporal Systems workshop, NeurIPS 2026
2025 —
2026
Which LLM Trading Behaviors Transfer
Working paper · ABClab, NTU CSIE · advised by Prof. Shih-Wei Liao

Designed an embargo-period protocol — all evaluation data dated after the teacher model's training cut-off — with multi-seed replication and matched-stack, label-shuffle, and threshold-sweep controls, to separate genuine model behavior from frontier-API response variance. Found that prompt-level differences for an LLM crypto-trading agent are mostly seed noise: none of four paired treatment effects had a 95% confidence interval excluding zero.

Fine-tuned Qwen3-8B with supervision-only LoRA on pre-embargo frontier-model decisions. One risk behavior — avoiding long exposure on crash days — survived replication and transferred to the small open model.

Experience

Jul —
Aug 2026
Trend Micro (TrendAI)
AI Engineer Intern · Taipei

Co-built ACE, an internal Automated Compliance Engine on the Claude Agent SDK that drives GitHub, AWS, and Playwright browser automation to collect ISMAP and ISO/IEC 27001 audit evidence. Implemented 40 ISMAP and 12 ISO controls end to end, a dashboard over a REST + WebSocket API, per-run cost monitoring, and multi-model routing to keep spend down.

Cut evidence collection for a control set from about five working days to under half a day. Internal teams tested it against real ISO evidence and reported edge cases that were fixed in turn.

Jul 2025
Shanghai Supercomputer Center
Testing Intern · Shanghai

Deployed a fully on-premise RAG knowledge base on an NVIDIA A800 80GB server — DeepSeek-R1-Distill-Qwen-14B served with vLLM, plus Qwen3-Embedding, Docling ingestion, and Open-WebUI, all containerized with no data leaving the local network. Ran functional, performance, and concurrent-load stress tests, then wrote the evaluation report: recommended vLLM over Ollama on throughput, security, and OpenAI-API compatibility.

Sep 2024 —
present
NASA — NTU CSIE Student System Administration
Team Leader · Taipei

Lead the student team that keeps the department's shared Linux workstations running — server operations, user support, environment maintenance, and Python / shell automation for routine tasks.

Selected projects

2026
figureout
Secondhand price-tracking platform for collectible figures in Taiwan

Real transaction data, trend analysis, and a Chrome extension for in-browser reporting.

Next.js · FastAPI · PostgreSQL · Docker
2026
tao_li
Delta-neutral funding-rate arbitrage on USDT-M perpetuals

Carry and near-settlement "snipe" modes, with dual-margin monitoring sized to survive ±30% moves on either leg independently.

Python · asyncio · Binance API
2025 —
2026
chao_bi
LLM-powered trading signal bot

Parses unstructured Telegram channel messages with a local Ollama model into structured Binance Futures orders — fully on-device inference, zero cloud. The project that pulled me into quantitative trading.

Python · Ollama
Aug 2026
AI Recommendation Diagnostics
2026 NTU "Build with AI" Hackathon

Diagnoses why AI assistants recommend a competitor's product, splitting the cause into information gap vs. product gap with a before/after validation loop.

FastAPI · AWS Bedrock

About

Senior in Computer Science & Information Engineering at National Taiwan University. Primarily a software engineer focused on LLM systems and AI infrastructure — building things that have to work, not just benchmark well.

My research interest is evaluation honesty: how to tell whether an AI system is actually doing the thing you think it is. Quantitative trading is the testbed I work in most, because markets give live, falsifiable feedback that a benchmark can't.

Currently looking ahead at graduate study and roles in LLM systems, AI infrastructure, or applied research. Open to collaboration.

Education
National Taiwan University — B.S., Computer Science and Information Engineering, Sep 2023–Jun 2027 (exp.). GPA 3.70/4.30 (3.85 in 2025–26).
Programming
Python, C/C++, Shell script; some TypeScript, Solidity
ML / LLM
LoRA / QLoRA fine-tuning, RAG, vLLM, PyTorch, Hugging Face, Ollama, Claude Agent SDK, evaluation-harness and prompt design
Systems
Linux administration, Docker, AWS, Playwright, Git, GPU / HPC clusters, networking
Languages
Mandarin (native), English (TOEIC 925), Japanese (JLPT N4 equivalent)
Honors
Participant, "Build with AI" 2026 NTU Hackathon, NTU Ventures — Aug 2026
Administrator, NTU CSIE Makerspace — Sep 2025–Jun 2026
Member, Information Department, NTU Student Association — Sep 2025–present

Get in touch

Email is fastest.