Python code quality: tools and aspects¶
Tuesday, June 2nd 2026, 19:00
Spektral, Lendkai 45, 8020 Graz
This time there will be a theme evening with several talks and a discussion on the overarching topic "Python code quality". While some say that more source code means more productivity, others say that every line of source code is a liability. In addition to the quantity, however, the characteristics of each line of source code are decisive for how well a software application behaves and how easy it is to make changes.
Although Python has a dynamic type system and thus a lot of information is only available at runtime, there are some tools to evaluate the maintainability, comprehensibility, efficiency, and robustness of an application purely on the basis of the source code.
During this Meetup, we'll take a look at some of them.
Pre-commit hooks and prek¶
Balasz
TBD
Type checking with Ty¶
Thomas
Slides and code examples
Python type hints are optional and never validated by the interpreter, but can still be helpful for documentation and navigating withing an IDE. Using type checkers, they can also be used to statically check for type violations.
This talk introduces the basics of type hints including a few more advanced topics like forward references and generators.
In conclusion, Python type hints can be helpful to create better code, but are still a moving target and have some idiosyncrasies.
Python Code Quality: Humans & AI Agents¶
Sebastian
Agentic Coding & "Agentic XP"¶
- Agentic Coding Levels:
- White Box: Only humans write code (critical systems) [1, 2].
- Black Box: AI writes code; humans treat it as a black box (small tools/PoCs) [1, 2].
- Grey Box: Partially AI-generated with shared human-AI understanding (mid-tier tasks) [1, 2].
- Resource: Stop software "slop" (Mario Zechner) [1, 2].
- Extreme Constraints Strategy: Uncle Bob Martin argues developers shouldn't read AI-generated code to truly save time. Instead, wrap agents in extreme constraints (tests, quality metrics) [2].
- Resources: Bob Martin's Post | Bob Martin's Talk [3].
- "Agentic XP" (Extreme Programming for Agents): Product Owners write specifications, developers act as supervisors, and agents implement code and tests [3].
- Note: TDD is unnecessary in the agent loop—it triples agent execution time without bringing measurable quality improvements [3].
- Resource: TDD in the Agent Loop (Martin Fowler) [3].
Evaluated Python QA Tool Stack¶
A curated toolset to automate quality gates for AI agents.
| Tool | Focus & Purpose | Key Highlights / Constraints | Links |
|---|---|---|---|
| Behave | Behavior-Driven Development (BDD) via Gherkin specs [4, 5]. | Renewed interest since AI can auto-generate Gherkin specs and step implementations [4, 5]. | Docs |
| pytest-crap | Calculates the CRAP score combining Cyclomatic Complexity (CC) and Coverage (cov) [6]. | Formula: $CRAP(m) = CC(m)^2 \times (1 - cov(m))^3 + CC(m)$ [6]. • < 5: Excellent• > 30: Critical (requires refactoring) [7]. |
PyPI |
| mutmut | Mutation testing to uncover untested code branches [7]. | Pros: Reliably finds untested paths [8]. Cons: Heavy resource usage; not suitable for agents as-is due to a lack of LLM-friendly feedback [8]. |
PyPI |
| Bandit | Static security scanner for common vulnerabilities (e.g., SQLi) [8]. | Pros: Blazingly fast, low false-positive rate [10]. Cons: Limited detection; fails to find logical issues like unsafe redirects [9, 10]. |
Plugins |
Other Recommendations: pytest (unit tests), ruff (fast linter with auto-fix), and ty (type checker) [4].
Data harvesting your git history¶
Dorian
A Microsoft paper (2005) claimed :
"churn-based metrics predicted defects more reliably than complexity metrics alone."
This is language independend, so lets find our what we can use to analyse a git source code repository ..
the 20 most churned files¶
see the most changed files.
who build the repository¶
see who commited how often.
who build it an the last 6 month¶
see who commited how often lately.
where were bugs documented¶
what does the log say about problems (examplary terms).
git log -i -E --grep="fix|bug|broken" --name-only --format='' | sort | uniq -c | sort -nr | head -20
when was alot of work done¶
activity in the project matters too.
when did the repo burn¶
see documented troubles (examplary terms)
Links¶
Here are some links mentioned during our discussion after the talks.
- PyGraz community chat: We have a Discord channel now. Due to its increasing enshittification, we also discussed a few possible open source alternatives.
- Video: 4 words triggered a war: The story of Mitchell Hashimoto's "controversial" quote
I read the code.
- Chatto, a possible open source alternative to Discord.