Code

Aggregating Keyword Categories for Lexer Tokenization

In a Pygments lexer, distinct subsets of lexical terms—such as core keywords, alphabetic operators, extra keywords, constructors, and block delimiters—can be categorized into separate multiline strings and then combined into a single master keyword list:

keywords += operators + extra_keywords + constructors + block_open + block_close keywords = keywords.split()

Invoking .split() converts the concatenated strings into a unified list of words, which can then be compiled into regex patterns or matched using Pygments' token classification rules.

0

1

Updated 2026-10-08

Tags

Prep Sessions

Tokenizing Source Code: Classifying Keywords and Operators in Pygments @ University of Michigan - Ann Arbor

Ch.1 Lexical Analysis - Tokenizing Source Code: Classifying Keywords and Operators in Pygments @ University of Michigan - Ann Arbor

Keyword and Operator Categorization - Tokenizing Source Code: Classifying Keywords and Operators in Pygments @ University of Michigan - Ann Arbor