AI Meets Science
A research lab for talented middle and high schoolers who test what AI tools actually do in science
In-person in Princeton, NJ · or hybrid from anywhere

The Opportunity
In this lab you measure what AI tools do in scientific work. You pick a tool and a real task in a field you care about, design a comparison that can come out either way, run it, and report what you find. The questions are about evidence: where does the tool hold up, where does it fail, and how would you know.
Tools now claim to read thousands of papers, generate hypotheses, write analysis code, and predict protein structures, and research institutions have adopted them quickly. Understanding of how well they perform lags behind that adoption, and the questions are open in a way that a careful, well-designed comparison can address.
AI Meets Science is one of the SoTS Research Labs, which follow the same methodology and format: see Princeton Labs for how semesters work and expectations from students participating in the program.
Recent Work in This Field
This field is active. Recent work compares AI tools against established methods in literature search, scientific code, statistical analysis, structure prediction, image analysis, and hypothesis generation.
Recent publications
Tools measured against established methods
- Elicit found 39.5% of the studies included in four evidence syntheses, against 94.5% for the original searches, though it surfaced some studies those searches missed (Lau & Golder, 2025).
- Asked for a numerical integration, a conjugate gradient solver and a parallel heat-equation solver in nine languages, ChatGPT produced Fortran and Rust solvers that needed fixing before they compiled (Diehl et al., 2024).
- SPSS and ChatGPT-4 returned identical results for t-tests and simple linear regression, but post-hoc analysis after one-way ANOVA showed discrepancies in mean differences and confidence intervals (Shahrul & Syed Mohamed, 2024).
- The accuracy of AlphaFold predictions varies, and they do not account for ligands, covalent modifications or other environmental factors (Terwilliger et al., 2023).
- Segmentation encoders pre-trained on over 100,000 microscopy images generalized better to micrographs taken under unfamiliar imaging and sample conditions than encoders pre-trained on ImageNet (Stuckner et al., 2022).
- Of 96 hypotheses GPT-4o generated for cardiotoxicity research, experts rated 13 as highly novel and 62 as moderately novel (Li et al., 2025).
Variety of AI Tools
The tools this lab works with are public, and most offer a free tier. Some require a paid subscription.
- Literature search: Elicit, Consensus, Scite, PaperQA
- Scientific code: GitHub Copilot, ChatGPT, Claude
- Data analysis: ChatGPT Advanced Data Analysis and other code interpreters
- Structure prediction: AlphaFold, ESMFold, RoseTTAFold
- AI scientist platforms: Google AI Co-Scientist, FutureHouse
Prerequisites
Open to high school, middle school, and home school students curious about AI and at least one area of science. No coding or AI experience is required.
Full requirements
- Willing to read into a real literature and design your own study, not follow pre-made assignments
- Commitment to weekly meetings and work between sessions
- Some tools require a paid subscription. Which ones depends on what you choose to evaluate, and the cost is not covered by SoTS Labs