The first PrivDev study connects data-type labels from a code scanner to a privacy vocabulary. That sounds straightforward until a label is ambiguous, the vocabulary has no close match, or a graph asserts more than the scanner could know. This article records what the first implementation accomplished and what later review changed. For the wider research programme, see the PrivDev introduction.
From a scanner label to a category#
Bearer CLI provides the 122 input data-type labels used in this study. The pipeline relates them to personal-data categories in the Data Privacy Vocabulary’s PD extension. The output is an RDF/Turtle graph that can be queried with SPARQL. It records a proposed interpretation of each scanner label, not a description of every real processing activity in which that label might occur.
The distinction matters for terms such as “Emails.” A program might handle contact addresses or the content of messages. “Employee Files” is broader still. A mapping can be plausible for a label and wrong for a particular code path. No vocabulary lookup can recover context the scanner never observed.
What was evaluated#
In the original Layer 1 study, 43 labels matched a preferred vocabulary label exactly; 79 required a less direct mapping. Retrieval and language-model suggestions helped with the latter, followed by structural checks and human annotation. The first graph covered all 122 input labels and contained 118 distinct policy resources because some labels converged on the same category.
The reported formative checks included SHACL validation, an ontology scan, five SPARQL questions and a nine-person annotation round. The annotators supplied 711 judgments. Raw agreement was 0.721, while Krippendorff’s alpha was 0.251; Gwet’s AC1 was 0.682 and weighted AC2 was 0.877. Those measures answer different questions, and their spread cautions against reading one headline number as proof of mapping correctness. Ten of the 79 non-trivial labels were flagged for review.
These figures describe the original study, not a claim that every resulting mapping is correct in a live system. Structural validation shows that the graph fits a specified form. It cannot establish the meaning of an ambiguous label or the legal position of a real organisation.
What the next iteration corrected#
A later review found that the original method could overstate what a data-type label established. It referred to specific GDPR articles where the processing context was unknown, and some identifiers did not exist in the pinned DPV vocabulary. The project changed the method to preserve the distinction between ordinary, special-category and criminal-offence data, to allow explicit abstention, and to present legal-basis schemes as candidates for review rather than conclusions.
The revised full run is recorded in the project as human-reviewed and structurally checked, but artifact promotion and wider evaluation remain separate steps. That difference matters: a completed experimental run is not a released compliance system. The wider security-to-privacy bridge, which would include software weaknesses and operational guidance, has not yet been evaluated end to end.
What this work can support#
The graph can make one part of a privacy review more inspectable. It lets a developer ask which concept was chosen for a scanner label and where that choice needs a second look. It cannot decide the processing purpose, legal basis, need for a DPIA or adequacy of a technical measure. Cryptography is a useful example: a scanner finding about cleartext storage can prompt a security investigation, but cannot by itself prove that a particular encryption measure is required or sufficient.
My main lesson from Layer 1 is that precision includes knowing when to stop. A formal graph and a fluent explanation can both create false confidence. The project is stronger when they show the evidence, the uncertainty and the point where a human decision is still needed.

