Projects

langxiao [dot] xie [at] berkeley [dot] edu LinkedIn GitHub

Atelier

A living map of a product, built out of the work coding agents do on it. An agent can touch dozens of files and report back in a single sentence, which leaves the person who owns the product steadily less able to say what it contains; Atelier reconstructs what exists, what changed in a session, and what changed because of the last request. In its first iteration at Bric&Brac; source is private.

Purl

The product I work on as a software engineer at Bric&Brac, across the backend foundations and the client applications. Source is private.

wiktextract — Spanish deixis tags

A Wiktionary dump parser that feeds kaikki.org and a range of downstream lexical projects. My patch repairs the Spanish demonstrative pronoun tables, whose spanning title header was missing from the inflection map, so every form came out carrying an error tag alongside the wrong proximal, medial, or distal deixis. Merged.

lm-evaluation-harness — EconLogicQA

EleutherAI's framework for few-shot evaluation of language models. I added EconLogicQA, whose 130 test questions each present four interconnected events from a business or supply-chain narrative and ask the model to order them by logical rather than chronological precedence — sequencing economic cause and effect, where the harness previously covered economics only through recall-based MMLU slices.

inspect_evals — PaperBench blacklist monitor

A collection of evals for Inspect AI. PaperBench's blacklist monitor stops an agent fetching the paper's reference implementation, but it matched only clone URLs that put host and path either side of a slash, so scp-like SSH syntax, where a colon separates the two, went undetected. My patch normalises that form before the check runs.

Sparse Autoencoder Playground

A sparse autoencoder written to be read in one sitting. It is trained first on synthetic data whose features are known, so that recovery can actually be checked, and then pointed at GPT-2 activations. Ships with a terminal walkthrough that draws its plots as coloured blocks, so the whole thing runs without a browser.

Rotation Momentum Strategy

A quantitative sector rotation system for A-share ETFs. Uses a three-factor momentum score — bias, slope, and efficiency — to rank Shenwan Level-1 sectors monthly and select the top-scoring ETF. Includes a trend filter on the CSI 300 index and automated WeChat push notifications via Server酱 on the last trading day of each month. Backtests at 13.9% CAGR against 8.2% for the CSI 300.

Vocabulum

A browser-based spaced-repetition vocabulary app for German, French, Spanish, and Latin. Uses a simplified SM-2 algorithm with four rating buttons and CEFR-leveled word lists sourced from kaikki.org and OpenSubtitles frequency data. No backend — runs entirely from static files.