EdTech Tool · Multi-Agent System · Faculty Support

What if you could red-team your syllabus before students do?

An AI stress tester that throws three common student cheating strategies at a college assignment, then generates a research-grounded vulnerability report with specific revision recommendations for faculty.

Role

Designer & Builder

Built for

College faculty

Stack

Python, Streamlit, Claude API, DeepSeek

Status

Functional prototype

The Problem

Faculty are designing assignments blind.

Generative AI changed the landscape of college assessment overnight, but most faculty have no way to know how vulnerable their assignments actually are before students submit them. The current options are bad: ban AI (unenforceable), run submissions through detectors after the fact (unreliable, biased against non-native English writers), or hope for the best.

None of those options help an instructor improve the assignment itself. The question isn't "did a student use AI?" The question is: "does this assignment produce strong evidence of learning, or can it be satisfied by someone who never engaged with the material?"

That's a design problem. And the tool I built treats it like one.

The Tool

Three AI agents attack your assignment. An advisor tells you what broke.

The instructor pastes in their assignment prompt, rubric, and course context. The tool runs three AI agents — each simulating a common student strategy for using AI to complete the work — then passes all three outputs to an advisor agent that produces a structured vulnerability assessment.

Input

Assignment + Rubric

Faculty paste their prompt and grading criteria

→

Red Team

3 Student Agents

Each one attempts the assignment with a different strategy

→

Analysis

Advisor Report

Vulnerability score, rubric gaps, revision recommendations

The output isn't a pass/fail. It's a faculty-facing report — a vulnerability score (1–10), a breakdown of what each agent exploited, the specific rubric criteria that were easy to game, and concrete, research-cited recommendations for making the assignment stronger.

The Red Team

Each agent tests a different failure mode.

The three agents aren't random — each one models a documented student behavior pattern, grounded in research on cognitive offloading, specification gaming, and reflective practice under AI.

Agent 01

The Quick Draft Student

Tests → baseline prompt complexity

Pastes the assignment and asks for a complete response. No rubric analysis, no refinement, no strategy. If this agent produces something gradable, the assignment's prompt alone doesn't require enough course-specific reasoning to resist one-shot AI completion.

Agent 02

The Rubric Optimizer

Tests → rubric specificity and gaming surface

Feeds rubric criteria directly into the prompt and structures the response to explicitly hit each one. If this agent scores well, the rubric may be measuring visible compliance signals — section headers, keyword mentions, citation counts — rather than actual evidence of learning.

Agent 03

The Reflective Voice Student

Tests → the verifiability of personal reflection

Makes the output sound personal, emotionally authentic, and grounded in first-person experience. If this agent's fake reflection is indistinguishable from real student work, the assignment may rely too heavily on written voice as evidence of lived experience — without anchoring it to verifiable process or local context.

Research Foundation

Every recommendation cites published research.

The advisor agent doesn't generate generic advice. It draws from a curated knowledge base of 8 pedagogical research frameworks — and is constrained to only cite sources from that base. No hallucinated citations. No vague "consider making it harder." Every recommendation maps to a specific framework and explains why it works.

Research frameworks in the knowledge base

Bloom's Revised Taxonomy — Anderson & Krathwohl, 2001

Authentic Assessment — Wiggins & McTighe, 2005

Process-Based Assessment — Yancey, 1996; Nilson, 2015

Metacognitive Reflection — Pintrich, 2002; Flavell, 1979

Specification Grading — Nilson, 2015

Situated Learning — Lave & Wenger, 1991

AI-Transparent Pedagogy — Mollick & Mollick, 2023

UNESCO AI in Education — UNESCO, 2023

This was a deliberate design constraint. Faculty trust research they can trace. An AI tool that generates unverifiable suggestions for how to resist AI isn't useful — it's ironic. The tool earns its credibility by being more rigorous about sourcing than the problem it's trying to solve.

Design Decisions

The choices that shaped the tool.

Constructive tone, not punitive.

The advisor is prompted to use language like "vulnerable when" and "stronger evidence of learning" — never "trivially fakeable" or "cheating." Faculty adopting AI tools shouldn't feel shamed for having assignments that were designed before generative AI existed. The tool is a design partner, not an auditor.

No AI detection. By design.

The tool deliberately avoids AI-detection scores. Research — including work from Stanford — has raised serious fairness concerns about GPT detectors, particularly false positives for non-native English writers. Instead of asking "was this written by AI?", the tool asks a better question: "does this assignment design produce strong evidence of learning regardless of what tools students use?"

Two models, different jobs.

The student agents run on DeepSeek — fast, cheap, and representative of what students actually have access to. The advisor agent runs on Claude — better at structured analysis, research synthesis, and maintaining a constrained output schema. Each model is chosen for the cognitive shape of the task, not as a brand preference. Same multi-model methodology that runs across all my AI work.

Institutional context awareness.

The tool can be configured with institutional context — course structures, existing touchpoints, required deliverables — so recommendations fit into what the instructor already has rather than adding new workload. A recommendation that says "add an oral defense" is more useful when it knows the course already has GSI check-ins it can repurpose.

The Report

What faculty actually receive.

The tool produces a structured vulnerability assessment — not a wall of text. Every section is designed to be actionable in a 15-minute syllabus revision session:

Vulnerability score (1–10) — how exposed the assignment is to AI completion, with a plain-language explanation of what the score means.

Agent performance breakdown — what each agent produced, an estimated rubric score for each, and one sentence on what the result reveals about the assignment's design.

Rubric gaps — specific criteria that could be satisfied without strong evidence of learning. Not "the rubric is weak." Instead: "Criterion 3 (demonstrates critical thinking) can be satisfied by restating course concepts without applying them to specific observations."

Research-backed recommendations — concrete rewrites, each citing a specific framework. Not "make it harder." Instead: "Reframe from 'analyze a social movement' to 'analyze how [specific organization from Week 7] applied theory in their 2023 campaign' (Wiggins, Authentic Assessment)."

The full report is downloadable as a text file — ready to share with a department chair, a teaching center, or a committee reviewing assessment policy.

Why It Matters

Shifting the conversation from detection to design.

For faculty

See your assignment through a student's AI lens — before they do.

Instead of discovering vulnerabilities after submissions come in, faculty can stress-test assignments during the design phase. The tool makes the invisible visible — showing exactly where and how an assignment can be completed without genuine engagement, and offering specific moves to close those gaps.

For institutions

A scalable alternative to the detection arms race.

AI detection is reactive, unreliable, and adversarial. Assignment redesign is proactive, equitable, and constructive. This tool supports the shift from policing AI use to designing assessments that produce evidence of learning regardless of what tools students bring.

For the field

Proof that product design thinking applies to education.

This isn't a research paper or a policy memo. It's a working product — designed, built, and deployable. The same design skills that shape a product's UX can shape how an institution responds to AI, and the tool is evidence of that translation.

The Argument

The best response to AI in the classroom isn't better detection. It's better assignment design.

This tool exists because the framing around AI in education has been wrong. The question was never "how do we catch students using AI?" The question is: "how do we design assessments that produce genuine evidence of learning in a world where AI is freely available?"

That's a design problem. I treated it like one — user research on the faculty experience, a multi-agent architecture grounded in pedagogical research, a report format designed for a 15-minute revision window, and an intentional choice to build toward improvement rather than surveillance.