Quick start

Install the package, fetch the data, run your first query

Install

R

install.packages("remotes")
remotes::install_github("ryanstraight/cybedtools")

Python

pip install cybedtools

Fetch the data and load a graph

cybedtools ships no framework text in the installed package. cybed_fetch() downloads and hash-verifies the public per-framework data release into a local cache (never re-downloading a file whose hash still matches), and load_graph() calls it internally and parses the result into one graph object. Both accept either the versioned framework slug (e.g. "nice-v2") or the short release-file slug (e.g. "nice"); see cybed_fetch()’s documentation for the two-vocabulary mapping. The data release backing the current package version is archived on Zenodo: 10.5281/zenodo.22884320.

library(cybedtools)

rdf <- load_graph()
from cybedtools import load_graph

graph = load_graph()

First queries

library(cybedtools)
library(dplyr)

# 1. What's in the corpus?
framework_summary |>
  select(
    framework_name,
    organizing_unit_count,
    elements_per_organizing_unit_with_examples
  ) |>
  arrange(desc(elements_per_organizing_unit_with_examples))
framework_name organizing_unit_count elements_per_organizing_unit_with_examples
DCWF v5.1 74 54.8
NICE v2.2.0 53 42.0
DigComp 3.0 26 34.0
ECSF v1 12 32.5
CyBOK v1.1.0 21 28.4
OTCCF v1.1 61 26.4
CCSSF 2022 59 22.8
CyQUAL 1.2.0 161 20.7
CSTA PK-12 CS (2026) 53 18.5
SCyWF 1.5 81 17.5
CSTA K-12 CS (Rev 2017) 25 10.3
SFIA 9 147 5.6
ACM/IEEE CSEC2017 8 5.0
Cyber.org K-12 v1.0 116 4.2
# 2. Which elements does an organizing unit carry?
unit_element_bindings(rdf) |>
  head(10)

# 3. Where do two frameworks say the same thing?
framework_similarity(rdf, from = "nice", to = "ecsf", n = 3)
from cybedtools import framework_summary, unit_element_bindings, framework_similarity

# 1. What's in the corpus?
framework_summary()[
    ["framework_name", "organizing_unit_count", "elements_per_organizing_unit_with_examples"]
].sort_values("elements_per_organizing_unit_with_examples", ascending=False)

# 2. Which elements does an organizing unit carry?
unit_element_bindings(graph).head(10)

# 3. Where do two frameworks say the same thing?
framework_similarity(graph, from_="nice", to="ecsf", n=3)

Next steps

Rebuilding the graph from source (maintainers)

Fetching the public release (above) is the path for using the corpus. Rebuilding it from primary sources is a separate, maintainer-only path, run from a clone of the repository: cybedtools does not redistribute framework source text, so each framework’s license governs how its source data is obtained and staged. docs/framework-data-sources.md documents the canonical retrieval URL, license, and SHA-256 reference checksum for each framework in the corpus.

For NICE and DCWF (US Government works, not subject to US copyright under 17 U.S.C. 105), the package ships pointers to the canonical NIST CPRT and DoD CIO downloads. For SFIA, ECSF, Cyber.org K-12, CSTA, CSEC2017, and DigComp, follow the steward’s redistribution policy as documented per framework. CyQUAL, CCSSF, OTCCF, SCyWF, and CyBOK are in the corpus on their stewards’ terms, several by written permission, and each carries its own terms. Read the per-framework entries before staging those.

scripts/000-build.R is the pipeline’s entry point: it runs every stage below in order.

# After staging source data under data/raw/<framework>/:
Rscript scripts/000-build.R

# Equivalently, stage by stage:
Rscript scripts/010-ingest-nice.R           # parse NICE -> tidy CSVs
Rscript scripts/010-ingest-dcwf.R           # ... and so on per framework
Rscript scripts/015-verify-ingestion.R      # cross-check ingestion output
Rscript scripts/020-assemble-jsonld.R       # uniform JSON-LD assembly
Rscript scripts/025-export-ntriples.R       # combined N-Triples graph
Rscript scripts/026-verify-graph.R          # graph invariants
Rscript scripts/030-export-release.R        # public per-framework data release
Rscript scripts/040-run-sparql.R            # canonical query outputs

The combined graph lands at data/processed/ntriples/_combined.nt, and the public release (what cybed_fetch() downloads) lands at data/processed/release/<version>/. Every query on this site reads from the combined graph.

Back to top