This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
What I Built
I built a two-part toolkit to solve the biggest time-sink in teaching large CS courses: grading massive, complex student projects. This isn't about auto-grading, which often fails for creative or multi-part assignments. Instead, it’s about automating the tedious, mechanical work so instructors can focus on providing meaningful feedback.
I built this for my friend, a CS professor who was spending entire weekends manually compiling, running, and debugging hundreds of student submissions for a single assignment.
The problem is simple: when you have 150 students submitting multi-file Java projects, just getting each one to compile and run with the correct test input is a logistical nightmare. It’s slow, error-prone, and mentally draining.
My solution is a two-part open-source toolkit:
CscGrader: A language-agnostic tool that automates the "collection of evidence" phase. It detects student projects, compiles them, runs them with predefined inputs, and captures all the output (stdout, stderr, exit codes, timing).
generateRubrics: A simple utility that takes a single master rubric file (Excel, PDF, etc.) and generates a personalized, ready-to-grade copy for every student on the roster.
Together, these tools transform a weekend of grunt work into a streamlined, semi-automated workflow, freeing up my friend to do what they do best: teaching.
Demo
The tools are designed for a command-line workflow. Here’s how they would be used together for a single assignment:
Step 1: Generate all the grading rubrics for the class.
bash
studentNames.txt contains a list of all students in the class
python3 generateRubrics.py -inputFile studentNames.txt -fileToCopy "Assignment-03-Rubric.xlsx" --assignment 03
This instantly creates a Rubrics/ folder with perfectly named files like JaneDoe-Assignment-03-Rubric.xlsx.
Step 2: Process all student submissions to collect evidence.
bash
The submissions folder contains one subfolder per student
The assignment profile tells CscGrader what to look for and how to run it
cscgrader batch ./submissions/ --assignment ./csc215-assignment-03.json --input test-input.txt
This processes every submission, compiles the Java code, runs it with the same test input, and saves a structured JSON file with all the results for each student.
Step 3: Grade.
My friend now opens the personalized rubric for a student in one window and the corresponding JSON evidence file in another. All the "did it compile? did it run? what was the output?" questions are already answered. They can now focus on code quality, design, and feedback.
Code
Here are the two repositories:
CscGrader: https://github.com/johnny603/CscGrader
generateRubrics: https://github.com/johnny603/generateRubrics
How I Built It
This project is a prime example of using open-source tools for practical, targeted automation.
Core Logic: Both tools are written in Python, relying only on the standard library. This was a deliberate choice to make them as portable and easy to run as possible for instructors who may not have admin rights or want to deal with virtual environments. The generateRubrics tool treats files as binary data, so it can copy .xlsx rubrics without needing libraries like openpyxl.
AI-Assisted Development: I used an open-weight model as a coding partner throughout the process. It was instrumental in designing the modular architecture of CscGrader, specifically the idea of using "assignment profiles" (JSON files) to make the tool reusable across different courses and languages.
Agent Harness: I used an open-source agent harness to automate the testing of CscGrader. The agent would create dummy Java projects with common errors (e.g., syntax errors, missing dependencies, ClassNotFoundException), run CscGrader on them, and verify that the tool correctly identified the failure category. This saved me hours of manual testing and made the tool far more robust.
The architecture of CscGrader is built around a modular pipeline: Detection → Compilation → Execution → Evidence. This is all orchestrated by a central runner, which uses language-specific "adapters" to handle the unique aspects of each programming language. The agent harness was key to building and validating this complex workflow.
Top comments (0)