From 2a09f6881e5340ef3fcff8fabafeb73409bbe3b8 Mon Sep 17 00:00:00 2001 From: Claude Date: Wed, 23 Sep 2026 13:26:11 +0000 Subject: [PATCH 1/3] Cache the parsed syntax dictionary per version getDictionaryByVersion() had no cache lookup: it built a fresh SQLSyntaxDictionary and SAX-parsed every grammar XML file on each call. A dictionaryVersionCache field already existed but was only ever read by getDictionaryByVersionAlt, an unused JDOM variant. Since CFMLParser's constructor calls it, every `new CFMLParser()` cost 13.3 ms of pure XML re-parsing. Measured over 200 constructions that drops to effectively zero. The cache key includes prefsSignature(fPrefs) so a later initDictionaries(prefs) with different preferences still gets a freshly loaded dictionary rather than a stale one. The loaded dictionary is immutable in practice -- the parser only reads it, and DictionaryManager already hands the same instance out through dictionariesCache -- so sharing it is consistent with how the rest of the class behaves. Scope note: this does NOT speed up a CFLint scan. CFLint builds one CFMLParser for an entire run, so it pays this once. It matters to a consumer that constructs a parser per file. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01GzpZFd4rnE1Yi2sVHAji35 --- .../main/java/cfml/dictionary/DictionaryManager.java | 11 +++++++++++ 1 file changed, 11 insertions(+) diff --git a/cfml.dictionary/src/main/java/cfml/dictionary/DictionaryManager.java b/cfml.dictionary/src/main/java/cfml/dictionary/DictionaryManager.java index 3430831..cf1cfed 100644 --- a/cfml.dictionary/src/main/java/cfml/dictionary/DictionaryManager.java +++ b/cfml.dictionary/src/main/java/cfml/dictionary/DictionaryManager.java @@ -357,6 +357,16 @@ public static SyntaxDictionary getDictionaryByVersion(String versionkey) { if (dictionaryConfig == null) throw new IllegalArgumentException("Problem loading dictionaryconfig.xml"); + // The parsed dictionary is immutable once loaded and every CFMLParser construction + // asks for the same version, so re-parsing the XML each time is pure waste. Key the + // cache on the preference signature as well so a later initDictionaries(prefs) with + // different preferences still gets a freshly loaded dictionary. + final String versionCacheKey = versionkey + "\u0000" + prefsSignature(fPrefs); + final SyntaxDictionary cachedDictionary = (SyntaxDictionary) dictionaryVersionCache.get(versionCacheKey); + if (cachedDictionary != null) { + return cachedDictionary; + } + // grab the cfml dictionary // Node n = dictionaryConfig.getElementById(CFDIC).getFirstChild(); Node versionNode = dictionaryConfig.getElementById(versionkey); @@ -400,6 +410,7 @@ public static SyntaxDictionary getDictionaryByVersion(String versionkey) { String filename = n.getAttributes().getNamedItem("location").getNodeValue().trim(); dic.loadDictionary(getDictionaryLocation(filename)); } + dictionaryVersionCache.put(versionCacheKey, dic); return dic; } From 61b20e74e6cd946a9eb9b37d93d35a27f9943f93 Mon Sep 17 00:00:00 2001 From: Claude Date: Wed, 23 Sep 2026 13:26:11 +0000 Subject: [PATCH 2/3] Hoist the repeated elem.toString() in visit() Jericho's Segment.toString() is source.subSequence(begin, end).toString() with no caching, so each call allocates a fresh String spanning the element's whole extent. For a that extent covers the entire body including nested content, and visit() called it four times in the cfif branch and twice in the cfset branch. Reading it once per branch cuts ~16% off a tag-heavy file: 33.2 ms to 27.8 ms over a 154 KB source of 60 blocks. Scope note: this only helps tag-based CFML. A scan of script-based components never reaches these branches -- parseCFExpression was called zero times over a 158-file .cfc corpus. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01GzpZFd4rnE1Yi2sVHAji35 --- .../src/main/java/cfml/parsing/CFMLParser.java | 12 +++++++----- 1 file changed, 7 insertions(+), 5 deletions(-) diff --git a/cfml.parsing/src/main/java/cfml/parsing/CFMLParser.java b/cfml.parsing/src/main/java/cfml/parsing/CFMLParser.java index aeb0be6..36cebfc 100644 --- a/cfml.parsing/src/main/java/cfml/parsing/CFMLParser.java +++ b/cfml.parsing/src/main/java/cfml/parsing/CFMLParser.java @@ -269,7 +269,8 @@ public void visit(final Element elem, final int level, CFMLVisitor visitor) thro // would otherwise be parsed as the expression "a = 1 /". // An expression can never legitimately end in '/' - division needs a right // operand - so removing a trailing one is safe. - String cfscript = elem.toString().substring(elem.getName().length() + 1, elem.toString().length() - 1) + final String elemText = elem.toString(); + String cfscript = elemText.substring(elem.getName().length() + 1, elemText.length() - 1) .trim(); if (cfscript.endsWith("/")) { cfscript = cfscript.substring(0, cfscript.length() - 1).trim(); @@ -295,21 +296,22 @@ public void visit(final Element elem, final int level, CFMLVisitor visitor) thro // added twice, reading past the end of the tag (and out of the string entirely once // the element sits far enough into the file). final int elemBegin = elem.getBegin(); - final int uglyNotPos = elem.toString().lastIndexOf("<>"); + final String elemText = elem.toString(); + final int uglyNotPos = elemText.lastIndexOf("<>"); int endPos = elem.getStartTag().getEnd() - 1 - elemBegin; if (uglyNotPos > 0) { - final int nextPos = elem.toString().indexOf(">", uglyNotPos + 2); + final int nextPos = elemText.indexOf(">", uglyNotPos + 2); // An unclosed or self-closing tag has no end tag; fall back to the element's own // extent rather than dereferencing null. - final int endTagBegin = elem.getEndTag() == null ? elem.toString().length() + final int endTagBegin = elem.getEndTag() == null ? elemText.length() : elem.getEndTag().getBegin() - elemBegin; if (nextPos > 0 && nextPos < endTagBegin) { endPos = nextPos; } } - final String cfscript = elem.toString().substring(elem.getName().length() + 1, endPos); + final String cfscript = elemText.substring(elem.getName().length() + 1, endPos); if (cfscript.length() > 0 && visitor.visitPreParseExpression("TAG", cfscript)) { final CFExpression expression = parseCFExpression(cfscript, visitor); From 7982d61d2831a4608d0ec43a3f57c34370554e59 Mon Sep 17 00:00:00 2001 From: Claude Date: Wed, 23 Sep 2026 13:26:20 +0000 Subject: [PATCH 3/3] Correct the CFLint parser-construction claim in CLAUDE.md CLAUDE.md stated that CFLint "constructs a new CFMLParser per file, so the expression cache is per-file and never accumulates across a scan." Both halves are wrong. CFLint.java has private CFMLParser cfmlParser = new CFMLParser(); as a field initialiser, so one parser serves the whole run -- an instrumented 158-file scan recorded exactly one construction. The expression cache therefore spans the entire scan. The claim is load-bearing: it invites optimising per-file construction cost, which is a pattern CFLint does not use. It is why the dictionary cache in this branch was expected to speed up a scan and does not. Replaces it with what a scan actually costs, measured rather than assumed: cold ANTLR DFA construction, 4676 ms on the first pass over 158 files against 86 ms on the second in the same JVM. That is fixed per-process cost, so 4x the files bought only 18% more wall clock. Also records that clearDFA() is process-global, since _decisionToDFA is static final on the generated parser -- calling it discards the warm-up for every parser in the JVM, not just the receiver. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01GzpZFd4rnE1Yi2sVHAji35 --- CLAUDE.md | 31 +++++++++++++++++++++++++++++-- 1 file changed, 29 insertions(+), 2 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index db2e1d0..6ea8bad 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -134,8 +134,35 @@ a large file. `cfmleditor/CFLint` is the main consumer. It declares the cfparser version **twice** — a `cfparser.version` property in `pom.xml` and a hardcoded coordinate in `build.gradle`. Bump both. -It constructs a new `CFMLParser` per file, so the expression cache is per-file and never -accumulates across a scan. +It constructs **one** `CFMLParser` for the whole run — `private CFMLParser cfmlParser = new +CFMLParser();` is a field initialiser on `CFLint`. A 158-file scan builds exactly one. So the +expression cache accumulates across the entire scan, not per file. + +(This paragraph previously said the opposite. Anything reasoning about per-file parser construction +cost — dictionary loading, cache locality — is reasoning about a pattern CFLint does not use.) + +### What a CFLint scan actually costs + +Almost all of it is ANTLR building its DFA the first time it meets each decision, not parsing. +Measured on 158 files in one JVM: + +| | | +|---|---| +| first pass (cold DFA) | 4676 ms | +| second and third pass (warm) | 86 ms | + +Same files, same process — **54×**. It is a fixed per-process cost, so scan time barely tracks +codebase size: 158 files took 5037 ms and 632 files took 5945 ms, a marginal ~2 ms per file. + +Two consequences worth keeping in mind: + +- Caching and allocation work inside cfparser will not move a CFLint CLI scan. Profiling one shows + ~91% of samples in ANTLR, `parseScriptBlock` accounting for 74% of wall clock, the SLL→LL fallback + firing zero times, and the expression LRU taking zero hits. +- A long-lived consumer (daemon, language server, IDE) amortises the warm-up completely and is the + single biggest available win. Conversely `clearDFA()` throws it away — and it is process-global + (`_decisionToDFA` is `static final` on the generated parser), so it clears every parser instance + in the JVM, not just the one it was called on. CFLint must stay on the same Java baseline. Its artifacts cannot load class file version 65 on an older JVM.