Patrick MacInnes

Most of the things I care about start with the same question: how does this work, and how do I know whether it works well?

I am a student in Grand Junction, Colorado, completing a double major in Computer Information Systems and Finance at Colorado Mesa University, with graduation expected in May 2027. That question shows up in school, in internship work at Tresic and Reiter USA, and in the projects I build on my own. The part of AI work I care about most is evaluation: defining what good means, measuring behavior, understanding uncertainty, and knowing when the evidence is strong enough to trust.

My LLM Judge Eval Harness is the clearest public example. It calibrates a quality threshold on purpose, measures held-out behavior, pays attention to bootstrap confidence-interval lower bounds, and fails closed when a promising score is not supported by enough evidence. I used the same methods as a production deploy gate at Tresic; that story is in An LLM Judge as a Deploy Gate. Other projects, from Nihongo Sensei to Webhook Lab, let me explore different kinds of systems end to end. Browse all of them on projects.

Outside school and work, I follow curiosity into subjects that do not obviously connect, including markets. I make Frenchcore remixes, return to anime, manga, and related stories, and have been a swimmer for most of my life. I want room for more than one thing without turning every interest into career optimization.

During school I am looking for remote internships and part-time work in AI evaluation, safeguards, agent quality, and related software engineering. After graduation (May 2027+) I want full-time roles in the same areas, especially in Denver, Southern California, or remote.

You can read about my work, browse projects, or email me.