Java (Swing) app to analyze GitHub repository activity for team projects (e.g., students).
🧭 Teaching focus: these metrics aim to approximate hands-on coding activity (especially in Java) in a learning context.
Some values may not match GitHub Insights → Contributors (see “📐 Metrics” and “❓ FAQ”).
- 🌳 Repository tree (left panel).
- 🧾 When you select a repo, shows repository-level metrics.
- 👥 Per-author table with commits/lines metrics (focused on
.java). - 🧩 Summary of file types in the default-branch snapshot.
- 🔄 Refresh from GitHub or work offline using a cache (
stats.dat).
- ☕ Java 17+ (required)
- 🗂️ Libraries included under
lib/for compiling (no Maven/Gradle required); the runnable JAR bundles them, so nothing extra is needed to run it - 🔑 A non-empty GitHub user field and token are required for an online refresh; use a read-only token with access to the repositories to analyze
On the first run the app creates a resources/ folder with default
config.properties and repositories.txt next to the app, and — if the required
user or token is not set — opens a Config dialog so you can fill in your GitHub user, token,
teacher account and the repository list from inside the app. You can reopen it any
time with the Config button.
You can also edit the two files by hand (see the Configuration section below); the Refresh from GitHub button re-reads them, so new tokens or repositories take effect without restarting. To prepare the files up front from the templates:
cp resources/config.properties.example resources/config.properties
cp resources/repositories.txt.example resources/repositories.txtRun the main class es.deusto.prog3.githubanalyzer.Main.
./build-jar.sh # Windows: build-jar.bat
java -jar github-analyzer-1.3.0.jarThis builds a self-contained versioned JAR (currently github-analyzer-1.3.0.jar): the application, the icons
and all third-party libraries are bundled inside — no lib/ folder is needed
to run it. On first run, the app creates a resources/ folder next to the JAR.
For an existing configuration, keep that folder with its config.properties and
repositories.txt next to the JAR; you can launch it from any directory (paths
resolve relative to the JAR's location). Those files stay
external and editable and are never bundled, since they hold your token and
student data.
# macOS / Linux (Windows: use ';' as the classpath separator)
javac --release 17 -cp "lib/*" -d bin $(find src -name "*.java")
java -cp "bin:lib/*" es.deusto.prog3.githubanalyzer.Mainstats.csv uses UTF-8 (with BOM) and ; as its separator so it opens cleanly in
common spreadsheet applications. Text fields are quoted, including values with
semicolons, quotes or line breaks. Values that could be interpreted as spreadsheet
formulas are exported as literal text.
./run-tests.sh # Windows: run-tests.bat (or "Run as > JUnit Test" in Eclipse)GitHub Actions
runs on every push and pull request targeting master. It verifies formatting,
runs the automated test suite with Java 17 and builds the runnable JAR.
github.user=_USERNAME_ # Required configuration field (not used for authentication)
github.token=_TOKEN_ # Required for online refresh; use a read-only token
update.from.github=_<yes|no>_ # yes: online refresh | no: offline (use stats.dat)
repositories.file=resources/repositories.txt # Repo list
stats.file=resources/stats.dat # Binary cache of last refresh
stats.csv=resources/stats.csv # CSV export
# (Optional) Exclude teacher account from expected-share calculations
teacher.user=_TEACHER_USERNAME_
teacher.email=_TEACHER_EMAIL_Custom paths for the repository list, binary cache and CSV export are supported. Saving user/token/teacher settings from the Config dialog preserves those paths.
One URL per line (GROUP-ID is optional):
https://github.com/OWNER/REPO1;GROUP-ID1
https://github.com/OWNER/REPO2;GROUP-ID2For private repositories and to reduce throttling, use a token with read access to the repos.
-
🧱 Total commits (unique)
Number of unique commits found (deduplicated by SHA), across analyzed history and known branches.
✅ Global activity indicator.
⚠️ Not the same as “Java commits per person” (see the table). Hover the label to see the breakdown: total = merge commits + non-merge commits, and the per-person Java commits only count non-merge commits that changed.java.Why the numbers don’t “add up” — Example: a repo shows 150 total commits. If 12 are merges, there are 138 non-merge commits. The per-person Java commits sum to some value ≤ 138 (e.g. 95), because commits that touched only non-
.javafiles (docs, configs, resources) or that were merges are excluded from the Java stats. -
📅 Creation date
Repository creation date from GitHub. -
⏱️ First commit / Last commit
Earliest/latest commit dates found in analyzed history.
📌 Useful to estimate real working period and detect end-of-period spikes. -
📄 Java Lines-Of-Code (LOC) (snapshot)
Current lines in.javafiles on GitHub's default branch (snapshot). 📌 Measures size, not effort. -
🔁 Java churn (added + deleted)
Total added+deleted lines in.javafrom analyzed commits (non-merge only).
✅ Proxy for editing effort.
⚠️ Can be inflated by formatting, generated code, or large pastes. -
🔗 External references
Occurrences of the standalone markersIAGorFUENTE-EXTERNAinside.javafiles on the default branch (matched as whole words, so they are not counted inside identifiers likeDIAGNOSTIC).
📌 Useful as a “reference/AI mention” signal (not proof).
Metric scope: File-based metrics (Java LOC, file types and external references) are a snapshot of GitHub's default branch only. Commit-history metrics traverse all known branches, deduplicate commits by SHA, and therefore can include work that is not present in the current default branch.
Identity: uses GitHub login when available; otherwise derived from commit author info.
Some identities may be merged (e.g.,noreply, same email local-part) to reduce duplicates.
Teaching-oriented identity matching: Students who are new to Git often make commits from different computers or IDEs without configuring the same name and email. To avoid splitting one student's work across several rows, the analyzer deliberately merges likely identities using GitHub login, email local-part and, when supported by another signal, normalized name. This prioritizes recovering a student's full contribution over strict Git-identity matching. In the uncommon case of students with coincident details, review the commit history before drawing conclusions.
-
✅ Java commits
Non-merge commits that touched at least one.javafile. -
➕ Java added / ➖ Java deleted
Added/deleted lines in.java(sum over non-merge commits). -
🔁 Java churn
added + deletedin.java.
✅ Better than “added only” because it accounts for refactors and fixes. -
📊 % Java churn
The user's share of the team churn — the total Java churn of the active contributors, excluding the teacher account.
Team shares sum to 100%. The teacher row shows-(excluded from the share model). -
🧩 Java files
Number of distinct.javafiles modified by the user. -
📅 First / Last commit (user)
First/last date of non-merge commit considered for that user.
The GUI adds a quick interpretation per person based on two concepts:
-
Churn share: the person's share of the team churn (active contributors, excluding the teacher)
[ share = \frac{userChurn}{teamChurn} ] whereteamChurnis the total Java churn of the active contributors (excluding the teacher). -
Expected share: what a balanced split would look like among active contributors (excluding the teacher account)
[ expected = \frac{1}{n} ] wherenis the number of “real” contributors:userChurn > 0ORjavaCommits > 0, excluding teacher.
Both
shareandexpectedare computed over the same population (active contributors, excluding the teacher), so they are directly comparable and the team shares add up to 100%.
Important: these are indicators, not an automatic grading system.
Use them to guide review, interviews, and code defense.
Badges are the main interpretation result. They appear:
- in the username cell (large marker),
- in the long tooltip,
- and in the status bar.
The GUI displays each icon at 24×24 px. The PNG source images are packaged with the application and scaled consistently by the GUI, so their appearance does not depend on emoji support or on a font installed by the operating system.
With expected = 1/n and share = userChurn/teamChurn (both over the same
population — active contributors excluding the teacher):
veryLow = expected * 0.5okMin = expected * 0.8okMax = expected * 1.2
Final badge (contiguous ranges, no gaps):
- TEACHER if the user matches
teacher.userorteacher.email - VERY_LOW if
userChurn == 0ORjavaCommits == 0ORshare < veryLow - BELOW if
veryLow <= share < okMin - BALANCED if
okMin <= share <= okMax - HIGH if
share > okMax
Flags do not change the badge: they are extra alerts to inspect patterns.
- Triggers when:
churnPerCommit >= repoAvgChurnPerCommit * 2.5- (with
javaCommits > 0andrepoAvgChurnPerCommit > 0)
- Suggests: large bursts per commit (mass paste / AI / generated code).
You hover a row and see:
- Username cell:
alice - Status bar:
Balanced contribution - Tooltip (long): shows expected share and no flags
Interpretation:
- Alice’s churn share is within ~±20% of the expected share for the team size.
- No additional signals triggered.
Row shows:
Interpretation:
- Bob contributes, but the Java churn share is below the expected range.
- Next step: check whether Bob contributed mainly in non-Java files, worked through PR reviews, or had a different role.
Row shows:
Interpretation:
- Carol has no Java churn or no Java commits, or a share well below half the expected one.
- Next step: confirm with the code defense — she may have worked only on non-Java parts, or contributed little.
Row shows:
Interpretation:
- Dave’s Java churn share is clearly above the expected share for the team size.
- This can mean a strong role — or an imbalance worth discussing with the team.
Row shows:
Interpretation:
- Eva’s total share might be modest, but her commits have unusually high churn per commit.
- LOC (Lines of Code): Number of lines in files (here,
.java) on the default branch. This is a size snapshot, not effort. - Churn:
added + deletedlines. Used as a proxy for “how much code was edited”. - Share (churn share): A user’s churn divided by the team churn (active contributors, excluding the teacher). Team shares sum to 100%; used to compare relative contribution within a team.
- Expected share:
1/n, wherenis the number of active contributors (excluding teacher). A baseline for “balanced” teams. - Non-merge commit: A commit that is not a merge commit. This app focuses on non-merge commits to better approximate authored edits.
- Deletion ratio:
deleted / (added + deleted). High ratios often indicate refactoring or cleanup. - Churn per commit:
userChurn / javaCommits. Spikes may indicate large pastes or generated code. - End-loaded work: Activity concentrated near the end of the project period (late commits).
GitHub uses its own heuristics (default branch, merges, attribution rules, etc.).
This app uses a teaching-oriented definition:
- “Java commits” = non-merge commits touching
.java - “Java churn” = added+deleted in
.javafrom non-merge commits - “Unique repo commits” = deduplicated by SHA
Goal: consistency for teaching interpretation, not to replicate GitHub UI.
Not necessarily. They may have:
- contributed mostly in non-Java files
- merge commits excluded from the Java stats
No. It only flags patterns compatible with large pastes/AI/templates.
Always confirm with code defense and understanding questions.
GitHub enforces API rate limits (roughly 5,000 requests/hour with a token, and
only ~60/hour without one). Analyzing many repositories — or repos with lots of
commits/branches — can hit that limit. The app requires a configured token for any
online refresh, and that token also needs access when repositories are private.
When fewer repos come back than configured, the app warns you; retry later, use a
token, or work offline with the cached stats.dat. Confirmed repositories are saved
progressively; if one repository fails during a refresh, its previous cached entry is
kept instead of being erased.
- ✅ Be transparent about what is measured and what isn’t.
- ✅ Use metrics to guide review—not as automatic grading.
- ✅ Consider roles and context (setup, reviews, non-Java work).
- ✅ Allow explanation and additional evidence.
- ✅ Avoid sharing identifiable metrics publicly.
- ✅ Keep
stats.datonly as long as necessary.
- Does not measure code quality (correctness, design, style).
- Commit habits differ (many small commits vs few large ones).
- Icons: Pixel perfect (Flaticon) — shown in the app footer.
The Java source code, build scripts and runnable JAR are licensed under the Apache License 2.0. Its attribution notice is available in NOTICE.
Documentation authored for this project is licensed under CC BY 4.0. Third-party libraries and icons are excluded from those grants and retain their own terms; see THIRD_PARTY_NOTICES.md.
For academic or teaching use, GitHub can generate the recommended citation from CITATION.cff through the repository's Cite this repository option.
Faculty of Engineering, University of Deusto — Academic year 2026-27.
The initial version of this codebase was developed in 2024 with partial assistance from ChatGPT (OpenAI).
From July to September 2026, the codebase was reviewed and audited using Claude Code (Anthropic) and Codex (OpenAI).
The resulting version was reviewed, tested, and refined to identify and correct issues within the scope of the performed verification activities.