Aggregating Keyword Categories for Lexer Tokenization
In a Pygments lexer, distinct subsets of lexical terms—such as core keywords, alphabetic operators, extra keywords, constructors, and block delimiters—can be categorized into separate multiline strings and then combined into a single master keyword list:
keywords += operators + extra_keywords + constructors + block_open + block_close keywords = keywords.split()
Invoking .split() converts the concatenated strings into a unified list of words, which can then be compiled into regex patterns or matched using Pygments' token classification rules.
0
1
Tags
Prep Sessions
Tokenizing Source Code: Classifying Keywords and Operators in Pygments @ University of Michigan - Ann Arbor
Ch.1 Lexical Analysis - Tokenizing Source Code: Classifying Keywords and Operators in Pygments @ University of Michigan - Ann Arbor
Keyword and Operator Categorization - Tokenizing Source Code: Classifying Keywords and Operators in Pygments @ University of Michigan - Ann Arbor