↓Skip to main content
  1. Digital Odyssey/

PrivDev: What a Code Scanner Can Tell Us About Privacy

·871 words·5 min·
Digital Odyssey Privacy Software-Engineering Research Cryptography Knowledge-Graphs
Simon Bernbeck
Author
Simon Bernbeck
Software engineer and master’s student in Informatics at PUC Rio de Janeiro, originally from Germany. Writing about what I learn, where I go, and what I find worth passing on.
Table of contents

A developer can search a codebase for personal data and still miss the question that matters: what should the team do with the finding? A scanner may report an email address or a storage pattern. It cannot see the full purpose of processing, the people involved, every technical safeguard, or the organisation’s legal decisions. PrivDev began with that gap between a technical observation and a responsible privacy review.

PrivDev is my research project at PUC-Rio’s AISE Laboratory. Its working idea is a security-to-privacy bridge: preserve what a scanner actually observed, connect it to a shared vocabulary, and make the reasoning available for inspection. I want developers to receive a useful starting point while leaving decisions that require context with the people who have it.

This article introduces the research programme. A separate article on Layer 1 looks more closely at the first mapping study.

The observation is smaller than the decision
#

Suppose a scanner identifies code that handles a field called email. That is evidence about the analysed code. It does not establish whether the field contains a person’s contact address, message content, a test fixture, or something else. Even if the value is personal data, the field name alone says nothing conclusive about why it is used or which legal basis applies.

The same distinction matters for cryptography. A tool might see a value written to browser storage without an encryption operation on the path it analysed. That is a reason to investigate. It is not proof that the value is unprotected everywhere: protection may sit in another layer, or the better answer may be to avoid storing the value at all. Encryption is one possible technical measure, not a universal legal conclusion generated by a scanner finding.

I use three questions to keep the bridge honest:

  1. What was observed in the code or in the scanner’s output?
  2. Which privacy concept is a defensible interpretation of that observation?
  3. What additional facts would a developer, security specialist, or privacy professional need before acting?

An answer becomes more useful when those three parts remain distinguishable. A fluent explanation that conceals the gap between them can be worse than a terse warning.

The first bridge: data types to a common vocabulary
#

The implemented first layer starts with the 122 data-type labels in Bearer CLI’s taxonomy. It relates them to personal-data concepts in the Data Privacy Vocabulary’s PD extension, or DPV-PD. The result is a queryable graph. A reader can inspect the category selected for a scanner label and the evidence or review state behind it.

Some labels match vocabulary terms directly. Others require a judgment about meaning and breadth. The workflow combines retrieval, language-model suggestions, deterministic checks, and human review. Its later iteration became deliberately more cautious: it preserves distinctions between ordinary, special-category, and criminal-offence data; it can abstain when the evidence is insufficient; and it no longer presents a specific GDPR article as though a data-type name established its applicability.

That revision taught me something useful. A deterministic rule can still be wrong if its premise is wrong. Earlier graph output contained references to legal concepts that were either too specific for the available evidence or absent from the pinned vocabulary. The project records those problems and corrects its mapping method. The point of a knowledge graph is not to make a claim look formal; it is to expose the claim so it can be challenged.

What comes after a category
#

The wider research plan asks how to connect other scanner findings, including software weaknesses, to privacy risks and possible measures. A cryptographic weakness is a good example because it invites an easy overclaim: “no encryption found” can turn into “encryption is legally required here” in one careless step. That step needs a threat model, an understanding of the data and its use, and expert review. PrivDev’s proposed later layers are meant to preserve those boundaries, not erase them.

I also study how this guidance could fit into AI-assisted coding. An assistant can produce code quickly, but speed does not give it the missing processing context. A useful assistant should point to the exact finding, name its uncertainty, and ask for the facts that would change the decision. It should be able to say “I cannot establish that from this scan.”

This is a research direction rather than a finished compliance product. The implemented and evaluated work is the first mapping layer. The broader bridge and its behaviour in a real development workflow still need to be built and tested. Even a structurally valid graph is not evidence that a particular system complies with the GDPR.

What I hope to learn
#

I want to know whether a traceable bridge helps developers ask better questions earlier: what data is involved, which safeguards exist, who can confirm the purpose, and where does a specialist need to intervene? Those are testable questions for later evaluation. For now, PrivDev offers a disciplined way to examine one narrow part of the problem and a record of where that approach has already needed correction.

The work is most useful when it makes uncertainty visible. Code can supply evidence. Vocabularies can supply shared terms. Neither can replace the context and accountability behind a privacy decision.

Related articles

PrivDev: Why a Cryptography Finding Needs Context
·698 words·4 min
Digital Odyssey PrivDev Cryptography Privacy Security Research
Doseframe: Making Perfusor Calculations Easier to Check
·793 words·4 min
Digital Odyssey Doseframe Healthcare Software Software-Engineering Research Patient Safety
PrivDev: Turning Static-Analysis Findings into Privacy Knowledge
··621 words·3 min
Digital Odyssey Privacy AI Software-Engineering RAG Knowledge-Graphs GDPR Research
Doseframe: A Calculation Is Only as Trustworthy as Its Inputs
·906 words·5 min
Digital Odyssey Doseframe Healthcare Software Patient Safety Software Architecture Usability
Project Planning Pipeline: Plan in Obsidian, Assisted by AI
··1746 words·9 min
Digital Odyssey Obsidian AI Project-Planning CLI Productivity Knowledge-Management Guide