How it works

read -> parse JSON -> choose @context -> skolemize -> JSON-LD to RDF -> SHACL -> humanise

The middle of that pipeline is standard JSON-LD processing and a standard SHACL engine (rdf-validate-shacl). The work of this library is at the two ends.

The problem

SHACL validates RDF. RDF has no notion of "line 10" or of address[0].postcode. It has triples and nodes, and most nodes in a converted JSON-LD document are blank nodes with machine-generated labels. A conformant engine will faithfully report:

Violation  path=https://example.org/postcode  focus=_:b3
           "Value does not match pattern ^[A-Z]{2}[0-9]$"

Everything about that is correct, and nobody can act on it.

Skolemization: putting results back where they came from

jsonld.toRDF mints its own blank-node labels, and those correspond to nothing the user wrote. So nothing is allowed to stay anonymous. Before conversion, the document is walked and every node object without an @id is given one, such as urn:dsv:node:7. The walk records where each node was found: the JSON pointer, the dotted path, the enclosing @type, the nearest ancestor @id, and the line and column. Afterwards, every focus node in the SHACL report is either an @id the user wrote or an identifier that decodes straight back to address[0].

This is safe because it adds no triples. @id is a node's identity, not a property. Cardinality constraints, sh:closed and sh:ignoredProperties are all unaffected. The one thing it would break is a shape requiring sh:nodeKind sh:BlankNode. The loader checks for that every time and turns the mechanism off when one appears. The tests also validate documents both ways and assert that the verdicts are identical.

The synthetic identifiers are internal. The tests also assert that urn:dsv:node: never appears in any output format.

Choosing a context

A context supplied to the validator always wins. This is what lets one set of shapes check documents whose own @context points somewhere else, or nowhere. Without one, the document's own context is loaded. URLs and relative paths are inlined up front, so the naming index sees the whole context. They are cached, so a batch naming the same context downloads it once.

Saying what is wrong

SHACL engines produce generic default messages ("Less than 1 values"). Parsing those would be fragile and would still say nothing useful. Instead, the reporter goes back to the shape that raised each result. It reads sh:description, sh:minCount, sh:pattern, sh:in and the other parameters, and builds a sentence from them. If a shape's description says "Postcode in standard format (e.g. AB1 2CD)", the hint's example is lifted straight out of it. A shape with its own sh:message was written for people, so that message is used verbatim.

One mistake often trips several constraints at once. For example, a value outside a vocabulary fails both sh:in and sh:class. Only the more specific complaint is kept.