| .. | ||
| lib | ||
| test | ||
| inventar.mjs | ||
| README.md | ||
Aere crypto inventory
A cryptographic inventory of source code. It reads a directory, finds where cryptography is used (algorithms, key sizes, curves, modes, TLS settings, JWT algorithms, key and certificate files), classifies each use by its exposure to a quantum computer, suggests a migration target, and writes the result as a CBOM (Cryptography Bill of Materials) in CycloneDX 1.6 format.
It is a single Node.js script with no dependencies: only built-in node:* modules. Nothing from the
scanned tree is executed, and no network request is made.
node tools/crypto-inventory/inventar.mjs scan ./my-service --out cbom.json --summary
What it detects
Detection is pattern-based on the source text, after comments (and Python docstrings) are removed. An API name that appears only inside a comment or inside a string literal is not reported as a use.
| Language | What is recognized |
|---|---|
| JavaScript / TypeScript | node:crypto (createHash, createHmac, createCipheriv, createSign/createVerify, generateKeyPair(Sync) for rsa, rsa-pss, dsa, ec, ed25519, x25519, dh, ml-dsa-*, ml-kem-*, slh-dsa-*, createECDH, getDiffieHellman, publicEncrypt, pbkdf2, hkdf, crypto.sign/verify, createPrivateKey/createPublicKey, X509Certificate); WebCrypto subtle (RSA-OAEP, RSASSA-PKCS1-v1_5, RSA-PSS, ECDSA, ECDH, Ed25519, X25519, AES-GCM/CBC/CTR/KW, HMAC, HKDF, PBKDF2, digest); TLS options minVersion, maxVersion, secureProtocol, ecdhCurve; libraries from imports: ethers, web3, viem (secp256k1 ECDSA), secp256k1, elliptic, node-forge, jsonwebtoken/jose (JWS algorithms), @noble/curves, @noble/post-quantum, @noble/hashes, crypto-js, tweetnacl/libsodium, node-rsa, bcrypt/argon2 |
| Python | cryptography.hazmat (rsa, ec, dsa, dh, ed25519, x25519, padding, hashes, algorithms/modes, AEAD classes); PyCryptodome (Crypto/Cryptodome: RSA, DSA, ECC, ciphers and modes, PKCS1_OAEP, pss, DSS, hash .new); hashlib (including hashlib.new and pbkdf2_hmac); hmac with digestmod; ecdsa; PyJWT / python-jose algorithms; oqs (liboqs-python); ssl protocol constants and TLSVersion |
| Java | KeyPairGenerator, Signature, Cipher, KeyAgreement, MessageDigest, Mac, KeyGenerator, KeyFactory, SecretKeyFactory, KEM, SSLContext getInstance("..."); ECGenParameterSpec, NamedParameterSpec; setEnabledProtocols; JSSE system properties (jdk.tls.namedGroups, protocol lists); Bouncy Castle imports and post-quantum parameter sets (org.bouncycastle.pqc, MLKEMParameters, MLDSAParameters, SLHDSAParameters, FalconParameters, ...) |
| Go | imports of crypto/rsa, crypto/ecdsa, crypto/elliptic, crypto/ecdh, crypto/ed25519, crypto/md5, crypto/sha1, crypto/des, crypto/rc4, crypto/mlkem, crypto/tls, golang.org/x/crypto/..., go-ethereum crypto, CIRCL ML-KEM/ML-DSA; calls with parameters (rsa.GenerateKey size, ecdsa.GenerateKey curve, hmac.New hash, mlkem.GenerateKey768); tls.Config MinVersion, MaxVersion, CurvePreferences (including tls.X25519MLKEM768) |
| Any text file | complete PEM blocks: certificates (parsed for subject, validity and key algorithm), public keys, private keys |
| Manifests | package.json, requirements*.txt, pyproject.toml, pom.xml, build.gradle(.kts), go.mod: cryptographic libraries that are declared |
In Go an unused import does not compile, so an import of a crypto package is evidence of use; the
import-level finding is dropped when the same file has a call-level finding that covers it. A blank
import (_ "crypto/sha512") is reported with a note that it only registers an implementation.
When the algorithm is given by a variable, the tool resolves it only when the file assigns that name
exactly once, to a plain string literal, and the name is not a function parameter or loop variable
(template literals whose ${...} parts all resolve this way are also resolved). The finding then
records where the value came from (resolvedFrom). In every other case the finding is unknown: the
tool does not guess. At most 64 distinct variables are resolved per file; past that, variables in that
file are reported as unknown, and the file is counted (files-over-resolution-limit in the CBOM, a
line in the text summary).
Cost on hostile input
The scanner reads repositories it did not write, so its cost is kept linear in the size of each file,
also for input built to be slow. Version 0.1.0, which was never published, had six places where a
crafted file cost quadratic time, minutes for a single file of the default maximum size: PEM block
search, JavaScript regex literals, Python triple-quoted strings on a long line, variable resolution
repeated for every call, open-ended parameter-list patterns used by that resolution, and the
assigned-variable lookup in Java. An adversarial review found them on 2026-09-29 and they were
rewritten for 0.2.0; the rewritten forms give the same findings as the old ones: on 2026-09-29,
test/echivalenta-0.1.0.mjs (which takes 0.1.0 from git history) found no difference on about 4,500
source files of the repository where the tool is developed, including three large npm packages, nor
on 300,000 generated inputs to the lexer and the PEM search. test/proba-timp.mjs keeps the cost
linear, and test/control-negativ-timp.mjs shows that it turns red when a quadratic form comes back.
One cost remains bounded rather than linear: the arguments of a call are read for at most 4,000
characters, so a file made of thousands of unclosed calls is slower (about 10 seconds per megabyte on
the machine where it was measured) but still finishes.
Classification
Every finding carries one class, a written reason and, when action is needed, a recommendation.
| Class | Meaning |
|---|---|
quantum-vulnerable |
Broken by Shor's algorithm on a cryptographically relevant quantum computer: RSA, DSA, finite-field DH, ECDSA, ECDH, EdDSA and X25519/X448 on any classical curve, secp256k1, pairing-based schemes. Also TLS configurations that allow TLS 1.2, which has no standardized post-quantum key exchange. |
weak-now |
Unsafe against classical attackers today: MD5, SHA-1, DES, 3DES, RC4, RC2, 64-bit-block ciphers, ECB mode, SSL and TLS below 1.2, RSA/DH below 2048 bits, curves below 112-bit security, JWT none. When the algorithm is also quantum-vulnerable (for example RSA-1024, or an RSA signature over SHA-1), the finding says so (quantumVulnerable: true). |
quantum-safe |
ML-KEM (FIPS 203), ML-DSA (FIPS 204), SLH-DSA (FIPS 205), Falcon, hybrid key-exchange groups such as X25519MLKEM768, AES, ChaCha20-Poly1305, SHA-2, SHA-3, HMAC and KDFs over them. For AES-128 the reason notes that Grover's algorithm gives at most a quadratic speed-up (NIST category 1) and that AES-256 is preferred for long-lived data. |
unknown |
The algorithm cannot be determined statically: it comes from a parameter, from configuration, from a key loaded at run time, or from a library that offers many algorithms without a recognized call. |
Recommendations are concrete but do not pretend a migration is simple. Examples: key exchange in TLS moves to X25519MLKEM768; signatures move to ML-DSA-65 or a hybrid classical+ML-DSA signature (the X.509 composite signature format was still an IETF draft when the rules were written); secp256k1 in a wallet moves at the account level to an account with a post-quantum owner, which is a protocol change and not a library swap; for JWT the tool says plainly that no final post-quantum JWS standard exists yet. Statements about standards and runtime defaults reflect the time the rules were written; check the current status before acting on them.
A finding states which algorithm appears where. It does not judge intent: a TLS scanner that deliberately offers classical groups to test a server will be reported like any other client.
Output
CBOM (CycloneDX 1.6)
bomFormat: "CycloneDX",specVersion: "1.6",serialNumber(urn:uuid:...),version: 1.metadata.timestamp,metadata.tools.components(this tool and its version),metadata.component(the scanned directory),metadata.properties(scan statistics: files seen, analyzed, not read and why, excluded directories, limits).- One component of
type: "cryptographic-asset"per distinct asset, with a stablebom-refderived from its content:cryptoProperties.assetType:algorithm,protocol,certificateorrelated-crypto-material.algorithmProperties:primitive,parameterSetIdentifier,curve,mode,padding,cryptoFunctions,classicalSecurityLevel,nistQuantumSecurityLevel(0 for quantum-vulnerable and weak algorithms; omitted when not defined).protocolProperties:type: "tls",version.certificateProperties: subject, issuer, validity, format.relatedCryptoMaterialProperties:type(private-keyorpublic-key),format,size. The value of key material is never written.oidwhere it is standard and certain.evidence.occurrences[]:location(path relative to the scanned directory),line,symbol(the API),additionalContext(the matched call, not the source line).propertiesin theaere:crypto-inventory:namespace:classification,quantum-vulnerable,reason,recommendation,languages,evidence-kinds,resolved-from.
- Declared libraries from manifests are components of
type: "library"withpurl, the manifest location, andaere:crypto-inventory:declared-only = true. Declaration is not use: a library can be declared and never called, and the cryptographic-asset components are the evidence of use.
--deterministic derives the serial number from the content and omits the timestamp, so two runs on
the same tree produce identical files.
Text summary (--summary)
Files seen, analyzed per language, checked for PEM only, and not read, with the reason (binary, too large, over the file limit, unreadable), plus excluded directories and symlinks not followed. Nothing is skipped silently: the scan fails with an accounting error if the categories do not add up to the number of files seen. Then counts per class and per language, the most frequent assets, the files with quantum-vulnerable findings, and the declared libraries.
Command line
node inventar.mjs scan <dir> [options]
--out <file> write the CBOM to <file> (default: stdout, unless --summary is given)
--summary print the text summary
--fail-on <list> exit 1 when findings match: vulnerable, weak, unknown, any (comma-separated)
--max-file-bytes <n> skip and count files larger than n bytes (default 1048576)
--max-files <n> read at most n files; the rest are counted, not read (default 50000)
--exclude-dir <name> do not descend into directories with this name (repeatable)
--no-default-excludes also descend into .git, node_modules, __pycache__, .venv, venv, ...
--deterministic reproducible output (content-derived serial number, no timestamp)
--findings-json <file> write the raw findings, one object per occurrence
Exit codes: 0 success, 1 the --fail-on condition is met, 2 usage or scan error.
--fail-on vulnerable matches every finding a quantum computer would break, including weak-now
findings that are also quantum-vulnerable.
In CI
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with:
node-version: 24
- name: Cryptographic inventory
run: node tools/crypto-inventory/inventar.mjs scan . --out cbom.json --summary --fail-on vulnerable
- uses: actions/upload-artifact@v4
if: always()
with:
name: cbom
path: cbom.json
Start with --fail-on weak if the goal is to block known-broken algorithms first, and add
vulnerable once a migration plan exists; failing a build on every ECDSA use from day one usually
only teaches people to disable the check.
As a module
import { scan, buildCbom, renderSummary, checkFailOn, classify } from './inventar.mjs';
const result = scan('./src', { maxFileBytes: 1 << 20 });
const cbom = buildCbom(result, { deterministic: true });
console.log(renderSummary(result));
What it does NOT do
- It does not see cryptography inside compiled binaries (
.class,.jar,.so,.dll, executables) or inside transitive dependencies. It reads the source you point it at, and lists the libraries your manifests declare. - It does not find hand-written implementations (for example a local Keccak or AES routine), calls
made through wrappers or reflection, dynamically computed imports, or algorithms chosen in
configuration files at run time. Those appear as
unknownat best, or not at all. - It does not prove that a use is exploitable, reachable or security-relevant. MD5 used as a cache key is reported like MD5 used for signatures; the reason text tells you what to check.
- It is static pattern analysis. Supported languages: JavaScript/TypeScript, Python, Java and Go. Kotlin, Scala, C#, Rust, C and C++ files are only checked for PEM blocks. Files that are not UTF-8 text (including UTF-16) are treated as binary and counted as not read.
- It does not execute anything from the scanned tree, and it makes no network requests.
Tests
node test/proba.mjs # 37 probes: fixtures per language with the exact expected findings
node test/control-negativ.mjs # breaks the scanner in a copy, one rule at a time (18 mutations), and
# requires the matching probe to fail while the rest of the suite still runs
node test/proba-timp.mjs # 8 hostile inputs, each at two sizes: the cost must grow linearly
node test/control-negativ-timp.mjs # puts each quadratic form back and requires the cost probe to turn red
node test/echivalenta-0.1.0.mjs # same findings as 0.1.0 (needs the git history; skipped without it)
Fixtures mark each expected finding on its own line (EXPECT: <name> | <class>), and a probe fails on
a missing finding, an extra finding, a wrong line or a wrong class. Negative fixtures cover comments,
docstrings, strings that mention crypto APIs, a file without cryptography, algorithms from
unresolvable variables, a binary file and the size and count limits. The private-key probe generates
a key at run time (none is stored in the repository) and checks that no part of it appears in any
output.
Developed and tested on Node.js 24; older versions are not tested.