Building a record of Canada’s military export industry

During four months at Project Ploughshares, I created a pipeline making efficient use of AI agents, among other tools, to find and consolidate information about Canadian military exports from information online, into their 50+ year old Canadian Military Industry Database (CMID). I created an improved and secure dashboard to automate what had taken researchers a full day every week, while also finding more transactions, all for about a dollar a day.

Employer
Project Ploughshares
Role
Research and development assistant
Term
May – August 2026
Program
B.Comp. Computer Science (Co-op), University of Guelph

Introduction: how I connected with Project Ploughshares, and a bit about them

I met Kelsey Gallagher, a senior researcher at Project Ploughshares, in April 2025, where he was a guest speaker at Civic Tech Waterloo Region (CTWR), a volunteer group I was a part of (and at the time of writing still am). CTWR is based in Waterloo and aims to help non-profits and local organizations using tech and design.

He explained how one large part of his work was maintaining the Canadian Military Industry Database (CMID), a record of international arms contracts awarded to Canadian companies, to support his research. The overall goal of that research is to advocate for stronger government compliance with national and multilateral arms control regimes, including the Arms Trade Treaty.

Before my co-op over the summer, I volunteered as part of a team from CTWR, aiming to make an MVP for Project Ploughshares to use. We worked with Kelsey, and ended up with a working prototype of an open-source intelligence tool. Scrapers pulled arms-trade news from Google Alerts and one other source, an LLM extracted possible Canadian military export deals, and researchers approved or rejected them in a web app before they were recorded. It proved the idea, but it was disjoint from their current workflow, and lacked features, for example deduplication, consideration of the reliability of different sources, backups, and more. My co-op project rebuilt it for production. It connects to the researchers’ real database, matches names to their existing company and country codes, removes duplicates, controls LLM cost, and runs securely, with weekly cloud backups, along with several new features.

Goals and learning outcomes

Still to write. This is a named element in the guidelines and it is the only one missing. Roughly 250 words, in your own words:

What you set out to learn going in and why. Which goals you met. At least one you did not meet, and why. Which skills came from coursework and which you picked up on the job.

The work

This project worked in a few stages. First was just setting up, and designing the system. I already had a basic idea of what I was doing from working with Civic Tech and seeing Kelsey’s reaction to the MVP. So for the first couple of weeks I decided on a tech stack (FastAPI backend, plain HTML/CSS/JS frontend), and created a documentation and testing framework. I set up the GitHub repo, Docker and a CI/CD pipeline, and infrastructure for my coding agents. I wrote deploy and setup scripts for the server, which is currently running on an old laptop in my bedroom (part of keeping the cost down, and it’s pretty stable to be honest), a VPN for securely connecting to the server, and health check monitoring that emails me upon incidents.

The second stage is where I’ll group the main bulk of the development work, up until late July. This is where I was checking in with Kelsey, but was doing large amounts of the development on my own. I took inspiration from the Civic Tech project, but was building everything from scratch. I built an escalating system of checks, that basically boils down to:

  1. Set up sources. Google Alerts, RSS news feeds, company and government page scraping and crawling. This gives us a long queue of URLs to scrape then process.
  2. Scrape the articles, and do varying levels of filters, all relatively cheap but in increasing precision and cost. This narrows it down to relevant articles.
  3. Once we have a much smaller list of articles of interest (about 2% of our original list), we hand it to our custom researcher agent, which searches the web and the database of 76k+ past transactions, to identify the supplier, recipient, amount, and details about the transferred materials, among other fields. It also can apply estimates of commonly transferred items, mined from past deals and pre-approved by researchers. It deduplicates transactions, finds and saves corroborating sources, and routes the item.

This is how we are able to find transactions very accurately, and still while casting a wide net, checking about 200 articles each day, for $1 per day.

I mentioned that items coming out of the research agent get routed, and this is an important part of optimizing Kelsey’s time spent reviewing items. Items fell into one of a few buckets. Each bucket and decision is always visible to Kelsey to audit, with the ability to adjust or undo a decision. The buckets are as follows:

The third and final stage was the last month-ish of the co-op. By this point, the core features had been built into the product, and I had been meeting with Kelsey about twice weekly. During the last month, we were meeting nearly every day, for anywhere from 30 minutes to sometimes 2 hours. These were user testing meetings, which gave me really good feedback to refine the product to meet his workflow, refine the output of the agents, and find bugs. Also, this served to help Kelsey learn to use the new tool.

In the end… well, not really the end

By the end of the summer, the pipeline was reliably finding transactions. Some metrics on the research side:

And the development side:

Working with Project Ploughshares gave me an appreciation for working on a small team, where I had a lot of autonomy over my work, and was seeing everything end-to-end. Also, it was very satisfying working on something as rewarding as peace research.

Though that concludes my co-op, I continue to work with Ploughshares, maintaining the project, and adding features and improvements here and there.

Acknowledgements

Thanks to Kelsey Gallagher and everyone at Project Ploughshares for making this a very enjoyable co-op. Thansk to Dr. Luiza Antonie for being my academic supervisor for this project. And thanks to the Mitacs program, and the sponsors of Ploughshares, for funding my work.