Skip to content

Repository files navigation

ASN.1 Parser

JSR

ASN.1 text parser in TypeScript. To clarify: this is not a Basic Encoding Rules (BER), Distinguished Encoding Rules (DER) encoder / decoder, etc. If you are attempting to serialize or deserialize ASN.1 data, this is not the correct module for your purposes. This module parses the textual ASN.1 specifications themselves, according to the syntax defined in the freely available ITU-T Recommendations X.680, X.681, X.682, and X.683.

Documentation

Security

It is ill-advised to allow this to be used with untrusted inputs. There have been denial-of-service vulnerabilities found, and there are probably more yet to be uncovered.

Usage Example

This is a test, completely copied and pasted here to showcase the capabilies and usage of this module:

import { strict as assert, strictEqual as assertEquals } from 'node:assert';
import { lex, grok, normalize, parse, correct, AssignmentType, TypeType } from '../dist/index.mjs';
import { test } from 'node:test';

const AuthenticationFramework = `
AuthenticationFramework {joint-iso-itu-t ds(5) module(1) authenticationFramework(7) 8}
DEFINITIONS ::= BEGIN

-- EXPORTS All

IMPORTS

  ATTRIBUTE, DistinguishedName, MATCHING-RULE, Name, NAME-FORM, OBJECT-CLASS,
  RelativeDistinguishedName, SYNTAX-NAME, top
    FROM InformationFramework informationFramework ;

SIGNATURE ::= SEQUENCE {
  algorithmIdentifier  AlgorithmIdentifier{{SupportedAlgorithms}},
  signature            BIT STRING,
  ... }

SIGNED{ToBeSigned} ::= SEQUENCE {
  toBeSigned    ToBeSigned,
  COMPONENTS OF SIGNATURE,
  ... }
  
END`;

test('the README example works', () => {
  const text = AuthenticationFramework;
  const lexResults = Array.from(lex(text));
  const parseResults = parse(text, lexResults);
  const modules = grok(text, parseResults);
  const normalizedModules = normalize(modules);
  correct(normalizedModules);
  const afmod = normalizedModules[0];
  assertEquals(afmod.name, 'AuthenticationFramework');
  const sig = afmod.assignments.SIGNATURE;
  assert(sig.assignmentType === AssignmentType.TypeAssignment);
  assert(afmod.assignments.SIGNATURE.type.typeType === TypeType.SequenceType);
  /** @type {import('../dist/index.mjs').SetOrSequenceType} */
  const seq = afmod.assignments.SIGNATURE.type.type;
  assertEquals(seq.rootComponentTypeList1.length, 2);
  const [ comp1, comp2 ] = seq.rootComponentTypeList1 ?? [];

  assertEquals(comp1.namedType.identifier, 'algorithmIdentifier');
  assertEquals(comp1.namedType.type.typeType, TypeType.DefinedType);
  assertEquals(comp1.namedType.type.type.reference, 'AlgorithmIdentifier');
  assertEquals(comp1.optional, false);
  assertEquals(comp1.text, 'algorithmIdentifier  AlgorithmIdentifier{{SupportedAlgorithms}}');
  assertEquals(comp1.default, undefined);
  // The offset of the start of this component in characters into the original text.
  assertEquals(comp1.production.location.startIndex, 341);
  // The offset of the end of this component in characters into the original text.
  assertEquals(comp1.production.location.endIndex, 341 + 63);

  assertEquals(comp2.namedType.identifier, 'signature');
  assertEquals(comp2.namedType.type.typeType, TypeType.BitStringType);
  assertEquals(comp2.optional, false);
  assertEquals(comp2.text, 'signature            BIT STRING');
  assertEquals(comp2.default, undefined);
});

Parsing and Groking Individual Productions

You don't have to parse or grok entire files at a time. You can parse individual productions, too! All of the "sub-parsers" and grokers are exported via parserFor and grokerFor.

import { strictEqual as assertEquals } from 'node:assert';
import { grokerFor, parserFor } from '@wildboar/asn1-parser';

const text = "MyType ::= INTEGER";
const ps = parserFor.TypeAssignment.start(tokens, text);
const ctx = createGrokContext(text, ps.definedEnumItems);
const ta = grokerFor.TypeAssignment(ps.cst, ctx);
assertEquals(ta.identifier, "MyType");

If you are re-parsing a substring of the entire ASN.1 file:

  • Supply a Location object to the startloc parameter of lex(text, startloc), which will add that location's startIndex to the start and end offsets of all lexical tokens (Productions). It will also adjust the line and column number. That way, the offsets for those tokens correctly point to their offsets in the original text, not the substring you re-lexed.
  • In the GrokContext, set textStartsAtOffset to the same startIndex of the Location you used for startloc in lex(). This is needed because the groking functions take the substring, not the whole string, so you have to "undo" the offset correction you did in lex().

Sorry this API is so dumb. I have never written a lexer / parser before this, and I originally had no intention or even thoughts about this being able to support parsing substrings of ASN.1. This is so janky because I had to not break compatibility.

I would also caution you that substring parsing is not very deeply tested.

Error Handling

In addition to the built-in error subclasses, Error and SyntaxError, this package also throws three custom error classes:

  • ASN1SyntaxError: thrown when there is an objective syntax error
  • ASN1SemanticError: thrown when there is a semantic error, such as a mismatching type and value, or a type reference referring to an object class assignment instead of a type assignment.
  • ASN1ParserExpectationError: thrown when the lexer / parser / groker encounters some unexpected state. Think of it as an assertion failure. If this happens, it might be a bug. Please let me know about it!

Each of these can be associated with a Production via their production field. Since the Production has a location field, you can use this to ascribe a location in the document to the error. ASN1SemanticError and ASN1ParserExpectationError can also have associated module names and assignment identifiers.

XML Value Assignments

This module theoretically supports reading XML value assignments, but this was never tested at all. It is very plausible that it works poorly, if at all.

JSON Exports

If you want to use this parser to do the parsing of ASN.1, then hand off friendlier data to another program, you can export the lexed tokens, the concrete syntax tree (CST), and the abstract syntax tree (AST) as JSON. The types used for these data structures were chosen so that they could serialize to JSON for consumption externally.

Example of exporting at each stage:

const lexResults = Array.from(lex(text));
const parseResults = parse(text, lexResults);
const modules = grok(text, parseResults);
const normalizedModules = normalize(modules);
fs.writeFileSync("./lexical-tokens.json", JSON.stringify(lexResults));
fs.writeFileSync("./cst.json", JSON.stringify(parseResults.cst));
fs.writeFileSync("./ast.json", JSON.stringify(normalizedModules));

See [doc/example.ast.json] for an example of what one ASN.1 module in ./ast.json would look like for you.

See [doc/example.lex.json] for what ./lexical-tokens.json might look like for you.

See Usage for slightly better documentation about this.

Command Line

After this package is installed, npx asn1parser concatenates one or more ASN.1 files and prints JSON (or ok for check):

npx asn1parser lex module.asn1
npx asn1parser cst module.asn1
npx asn1parser ast a.asn1 b.asn1
npx asn1parser check module.asn1
npx asn1parser --pretty ast -o ast.json module.asn1

ast runs grok, correct, and normalize so the JSON is ready for other programs. -p / --pretty indents with tabs. -o / --output FILE writes to a file instead of stdout. --help / help and --version / version are also supported.

Without a local install, the equivalent one-shot command is npx @wildboar/asn1-parser.

Module System and Environment

This module is published as an ESM module exclusively. If you are still using CommonJS, it is time to get with the times and switch to ESM. This module is published on both npmjs.com and jsr.io.

This module is intentionally run-time agnostic. It works on Node.js, Deno, Bun, and it probably would work on QuickJS and in any browser.

This module has a single run-time dependency, which itself has no further dependencies.

Building

You can build this library by running npm run build. The outputs will all be in dist. dist/index.mjs is the entry point where all of the symbols that constitute the public API are exported.

Testing

If you have Node.js installed, you can test using npm run node-test. To test with Bun, use npm run bun-test. To test with Deno, use npm run deno-test. There is only one Deno test, whose purpose is to kind of "smoke test" that this works on Deno.

You can check if this module has any problems with JSR by running npx jsr publish --dry-run --allow-dirty.

AI / LLM Usage Statement

None of the code in this repository was written AI / LLMs, except a few tests that were previously written run using Jest were converted to using the built-in NodeJS test runner, and Cursor was used to add some missing type annotations.

See Also

  • Meerkat DSA, an X.500 directory server that uses ASN.1 that was compiled to TypeScript using this module.

To Do

  • Make dependency on dependency-graph optional
  • Could the lexer take a TNext to change behavior, such as by returning a syntax error?
  • Convert ProductionType to string constants (maybe other enums too) (will require major version bump)

About

ASN.1 parser (the textual IDL, not BER, DER, etc.) in TypeScript

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages