Performant codemods
JSSG already calls your function once per matching file, so cost isfiles × work_per_file. Anything that walks or parses the rest of the repo inside that callback is quadratic. That is true for source transforms and for read-only miners.
Each file callback runs in a fresh QuickJS runtime. Module-level variables do not survive from one file to the next. Cross-file coordination must use codemod:workflow state (or the filesystem), not a module cache.
Use this diagnostic before you rewrite queries:
Local vs project-wide
Do not
readdir or parse() other files on every invocation. Do not getState a single repo-sized blob on every file: getState clones and reconstitutes the full JSON value, so a large shared index is still quadratic. Distinguish “not built yet” (undefined) from “built empty” so workers do not rebuild forever.
For set-diff, a complete index built once is correct. A partial/lazy index that grows as files are visited will edit or emit from incomplete sets.
Token gates must discriminate
A prefilter that still matches most of the repo is a parse tax. Prefer syntax-shaped checks such asfeatureFlags\s*: or a literal isActive('…') call over includes('featureFlags') when that substring also appears in plumbing names. Skip node_modules, tests, and generated dirs in both workflow exclude and any disk walk.
Insights (mining)
Insights binds each dashboard query to one workflow step and loads that step’sjs-ast-grep script from the published package. Metric names must be statically visible in that script (useMetricAtom + increment)—not only in a different package file or helper import. This is the selected step’s script, not necessarily the package entry. For mining, return null.
CLI workflows may use multiple JSSG steps or a shared language step so one parse serves several analyses. That does not apply to Insights: keep the metric emit in the bound step’s script, with a once-built sharded index when analysis is project-wide.
See Metrics for the useMetricAtom API.
Verify wall clock before you ship
jssg test on a handful of fixtures proves correctness, not scale.
- Time a no-op (
return null) on a few thousand mostly inert files, using the same entry path you will ship (workflow vs direct script). - Time the real package on that same target:
- Prefer
codemod workflow run -w workflow.yaml --target <dir>for packages that rely ongetSelector, workflowinclude/exclude, oroptions.matches. Directcodemod jssg rundoes not load the workflow selector or those globs. - For a standalone script with no workflow filters,
codemod jssg run --language <lang>is enough (--language, not-l;-lis only valid forjssg test).
- Prefer
- Optionally time a deliberate per-file full-repo walk so you know what quadratic looks like.
Best Practices
Consistent Patterns
- Use
const codemod: Codemod<TSX> = async (root, options) => {}withexport default codemod - Import proper types - use
Codemod,GetSelector, and language-specific types fromcodemod:ast-grep - Handle edge cases - always check for null/undefined values and use try-catch for async operations
- Use early returns - skip processing when possible, especially for sharding
- Batch operations - collect edits before committing
- Export getSelector - provide a selector function for performance optimization
- Optimization strategies - prefer early returns, single traversals, and batching edits; use specific patterns to reduce backtracking and false positives
Enterprise Considerations
- Idempotent transforms: Ensure transforms can be run multiple times safely
- Progress reporting: Log progress for long-running migrations
- Error handling: Gracefully handle unexpected input
- Performance: Optimize for large codebases. Follow Performant codemods: no-op first, then token gates, then a once-per-scan sharded index.
- Coordination: Use state for multi-repo coordination