Skip to content

Python code quality: tools and aspects

Tuesday, June 2nd 2026, 19:00
Spektral, Lendkai 45, 8020 Graz

This time there will be a theme evening with several talks and a discussion on the overarching topic "Python code quality". While some say that more source code means more productivity, others say that every line of source code is a liability. In addition to the quantity, however, the characteristics of each line of source code are decisive for how well a software application behaves and how easy it is to make changes.

Although Python has a dynamic type system and thus a lot of information is only available at runtime, there are some tools to evaluate the maintainability, comprehensibility, efficiency, and robustness of an application purely on the basis of the source code.

During this Meetup, we'll take a look at some of them.

Pre-commit hooks and prek

Balasz

TBD

Type checking with Ty

Thomas
Slides and code examples

Python type hints are optional and never validated by the interpreter, but can still be helpful for documentation and navigating withing an IDE. Using type checkers, they can also be used to statically check for type violations.

This talk introduces the basics of type hints including a few more advanced topics like forward references and generators.

In conclusion, Python type hints can be helpful to create better code, but are still a moving target and have some idiosyncrasies.

Python Code Quality: Humans & AI Agents

Sebastian

Agentic Coding & "Agentic XP"

  • Agentic Coding Levels:
  • White Box: Only humans write code (critical systems) [1, 2].
  • Black Box: AI writes code; humans treat it as a black box (small tools/PoCs) [1, 2].
  • Grey Box: Partially AI-generated with shared human-AI understanding (mid-tier tasks) [1, 2].
  • Resource: Stop software "slop" (Mario Zechner) [1, 2].
  • Extreme Constraints Strategy: Uncle Bob Martin argues developers shouldn't read AI-generated code to truly save time. Instead, wrap agents in extreme constraints (tests, quality metrics) [2].
  • Resources: Bob Martin's Post | Bob Martin's Talk [3].
  • "Agentic XP" (Extreme Programming for Agents): Product Owners write specifications, developers act as supervisors, and agents implement code and tests [3].
  • Note: TDD is unnecessary in the agent loop—it triples agent execution time without bringing measurable quality improvements [3].
  • Resource: TDD in the Agent Loop (Martin Fowler) [3].

Evaluated Python QA Tool Stack

A curated toolset to automate quality gates for AI agents.

Tool Focus & Purpose Key Highlights / Constraints Links
Behave Behavior-Driven Development (BDD) via Gherkin specs [4, 5]. Renewed interest since AI can auto-generate Gherkin specs and step implementations [4, 5]. Docs
pytest-crap Calculates the CRAP score combining Cyclomatic Complexity (CC) and Coverage (cov) [6]. Formula: $CRAP(m) = CC(m)^2 \times (1 - cov(m))^3 + CC(m)$ [6].
• < 5: Excellent
• > 30: Critical (requires refactoring) [7].
PyPI
mutmut Mutation testing to uncover untested code branches [7]. Pros: Reliably finds untested paths [8].
Cons: Heavy resource usage; not suitable for agents as-is due to a lack of LLM-friendly feedback [8].
PyPI
Bandit Static security scanner for common vulnerabilities (e.g., SQLi) [8]. Pros: Blazingly fast, low false-positive rate [10].
Cons: Limited detection; fails to find logical issues like unsafe redirects [9, 10].
Plugins

Other Recommendations: pytest (unit tests), ruff (fast linter with auto-fix), and ty (type checker) [4].

Data harvesting your git history

Dorian

A Microsoft paper (2005) claimed :

"churn-based metrics predicted defects more reliably than complexity metrics alone."

This is language independend, so lets find our what we can use to analyse a git source code repository ..

the 20 most churned files

see the most changed files.

git log --format=format: --name-only --since="1 year ago" | sort | uniq -c | sort -nr | head -20

who build the repository

see who commited how often.

git shortlog -sn --no-merges

who build it an the last 6 month

see who commited how often lately.

git shortlog -sn --no-merges --since="6 months ago"

where were bugs documented

what does the log say about problems (examplary terms).

git log -i -E --grep="fix|bug|broken" --name-only --format='' | sort | uniq -c | sort -nr | head -20

when was alot of work done

activity in the project matters too.

git log --format='%ad' --date=format:'%Y-%m' | sort | uniq -c

when did the repo burn

see documented troubles (examplary terms)

git log --oneline --since="1 year ago" | grep -iE 'revert|hotfix|emergency|rollback'

Here are some links mentioned during our discussion after the talks.

Map

View Larger Map