Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion docs/de/platform/automations/concepts.md
Original file line number Diff line number Diff line change
Expand Up @@ -84,7 +84,7 @@ Ein Werkzeug ohne Ausgabeschema liefert unstrukturierte Ausgabe. Soll daraus str

Tale prüft das ganze Dokument, wenn du es speicherst, wenn du eine Version bereitstellst und wann immer ein Client `validate_automation` aufruft. Ein **Fehler** beschreibt etwas, das sicher scheitert, oder Code, der die Analysegrenzen überschreitet. Er verhindert Speichern wie Bereitstellen. Eine **Warnung** zeigt auf etwas, das scheitern kann oder nichts Nützliches tut. Sie verhindert weder Speichern noch Bereitstellen; du entscheidest selbst, ob du etwas änderst. Jedes Problem nennt seine Node und sein Feld und, in einem Template, einer Bedingung oder in Code, den genauen Ausdruck.

Damit die Prüfung zügig bleibt, gelten für jeden Ausdruck und jeden `transform`-Code Grenzen von 8192 UTF-16-Codeeinheiten, 512 JavaScript-Tokens und 64 Verschachtelungsebenen in der Syntax oder im Syntaxbaum. Leerraum um einen Template-Ausdruck zählt nicht zu seiner Größe; Leerraum im Transform-Code zählt mit. Reiner Text außerhalb von Templates ist kein Code. Diese Grenzen können zuvor gültigen Code ablehnen. Kürze ihn oder verteile die Arbeit auf mehrere Nodes, bevor du erneut speicherst oder bereitstellst.
Damit die Prüfung zügig bleibt, ist jeder Ausdruck auf 8192 UTF-16-Codeeinheiten und 512 JavaScript-Tokens begrenzt. Für `transform`-Code gelten 16384 Codeeinheiten und 4096 Tokens. Beide dürfen höchstens 64 Verschachtelungsebenen in der Syntax oder im Syntaxbaum enthalten. Leerraum um einen Template-Ausdruck zählt nicht zu seiner Größe; Leerraum im Transform-Code zählt mit. Reiner Text außerhalb von Templates ist kein Code. Diese Grenzen können zuvor gültigen Code ablehnen. Kürze ihn oder verteile die Arbeit auf mehrere Nodes, bevor du erneut speicherst oder bereitstellst.

### Referenzen und Namen {#checks-references}

Expand Down
2 changes: 1 addition & 1 deletion docs/en/platform/automations/concepts.md
Original file line number Diff line number Diff line change
Expand Up @@ -84,7 +84,7 @@ A tool without an output schema is unstructured. To turn its text into structure

Tale checks the whole document when you save it, when you deploy a version, and whenever a client calls `validate_automation`. An **error** describes a definite failure or code that exceeds the analysis limits, and it stops both saving and deploying. A **warning** points at something that can fail or does no useful work; it never stops a save or a deployment, so you decide whether to act on it. Each problem names its node and field and, inside a template, a condition, or code, the exact expression.

To keep checks responsive, each expression and each `transform` body is limited to 8192 UTF-16 code units, 512 JavaScript tokens and 64 levels of syntax or syntax-tree nesting. Whitespace around a template expression does not count toward its size; whitespace in a transform body does. Plain text outside templates is not code. These limits can reject previously valid code: shorten it or split the work across nodes before saving or deploying again.
To keep checks responsive, each expression is limited to 8192 UTF-16 code units and 512 JavaScript tokens. A `transform` body may contain up to 16384 code units and 4096 tokens. Both retain a limit of 64 levels of syntax or syntax-tree nesting. Whitespace around a template expression does not count toward its size; whitespace in a transform body does. Plain text outside templates is not code. These limits can reject previously valid code: shorten it or split the work across nodes before saving or deploying again.

### References and names {#checks-references}

Expand Down
2 changes: 1 addition & 1 deletion docs/fr/platform/automations/concepts.md
Original file line number Diff line number Diff line change
Expand Up @@ -84,7 +84,7 @@ Un outil sans schéma de sortie produit une sortie non structurée. Pour transfo

Tale vérifie le document entier quand tu l’enregistres, quand tu déploies une version et chaque fois qu’un client appelle `validate_automation`. Une **erreur** décrit un échec certain ou du code qui dépasse les limites d’analyse : elle empêche d’enregistrer comme de déployer. Un **avertissement** signale ce qui peut échouer ou ne sert à rien. Il n’empêche jamais d’enregistrer ni de déployer, c’est donc toi qui décides d’agir. Chaque problème nomme son nœud et son champ et, dans un template, une condition ou du code, l’expression exacte.

Pour que les vérifications restent réactives, chaque expression et chaque corps `transform` sont limités à 8192 unités de code UTF-16, 512 tokens JavaScript et 64 niveaux d’imbrication dans la syntaxe ou l’arbre syntaxique. Les espaces autour d’une expression de template ne comptent pas dans sa taille ; ceux du corps `transform` comptent. Le texte ordinaire hors des templates n’est pas du code. Ces limites peuvent refuser du code auparavant valide. Raccourcis-le ou répartis le travail entre plusieurs nœuds avant d’enregistrer ou de déployer à nouveau.
Pour que les vérifications restent réactives, chaque expression est limitée à 8192 unités de code UTF-16 et 512 tokens JavaScript. Un corps `transform` peut contenir jusqu’à 16384 unités de code et 4096 tokens. Dans les deux cas, la limite reste de 64 niveaux d’imbrication dans la syntaxe ou l’arbre syntaxique. Les espaces autour d’une expression de template ne comptent pas dans sa taille ; ceux du corps `transform` comptent. Le texte ordinaire hors des templates n’est pas du code. Ces limites peuvent refuser du code auparavant valide. Raccourcis-le ou répartis le travail entre plusieurs nœuds avant d’enregistrer ou de déployer à nouveau.

### Références et noms {#checks-references}

Expand Down
70 changes: 64 additions & 6 deletions services/platform/lib/engine/core/syntax/limits.test.ts
Original file line number Diff line number Diff line change
@@ -1,11 +1,14 @@
import { beforeEach, describe, expect, it } from 'vitest';

import { nodeVmRunner } from '../../runners/node-vm';
import { createSandboxExecRunner } from '../../runners/sandbox-exec';
import { setCodeRunner } from '../runner';
import { evalCondition, evalTemplates, runCode } from '../template';
import type { Automation } from '../types';
import { validate } from '../validate';
import {
MAX_BODY_PARSE_TOKENS,
MAX_BODY_SOURCE_SIZE,
MAX_PARSE_DEPTH,
MAX_PARSE_TOKENS,
MAX_SOURCE_SIZE,
Expand All @@ -30,9 +33,6 @@ describe('parser analysis boundaries', () => {
'accepts %i source code units',
(size) => {
expect(expression(literal(size))).toMatchObject({ ok: true });
expect(parseBody(`return ${literal(size - 8)};`)).toMatchObject({
ok: true,
});
},
);

Expand All @@ -41,7 +41,9 @@ describe('parser analysis boundaries', () => {
ok: false,
limited: true,
});
expect(parseBody(`return ${literal(MAX_SOURCE_SIZE - 7)};`)).toMatchObject({
expect(
parseBody(`return ${literal(MAX_BODY_SOURCE_SIZE - 7)};`),
).toMatchObject({
ok: false,
limited: true,
});
Expand All @@ -54,7 +56,7 @@ describe('parser analysis boundaries', () => {
).toMatchObject({ ok: true });
expect(
parseBody(
`return ${'!'.repeat(extra)}${array((MAX_PARSE_TOKENS - 4) / 2)};`,
`return ${'!'.repeat(extra)}${array((MAX_BODY_PARSE_TOKENS - 4) / 2)};`,
),
).toMatchObject({ ok: true });
});
Expand All @@ -65,10 +67,31 @@ describe('parser analysis boundaries', () => {
limited: true,
});
expect(
parseBody(`return !!${array((MAX_PARSE_TOKENS - 4) / 2)};`),
parseBody(`return !!${array((MAX_BODY_PARSE_TOKENS - 4) / 2)};`),
).toMatchObject({ ok: false, limited: true });
});

it.each([MAX_BODY_SOURCE_SIZE - 1, MAX_BODY_SOURCE_SIZE])(
'accepts %i transform code units without enlarging expressions',
(size) => {
expect(parseBody(`return ${literal(size - 8)};`)).toMatchObject({
ok: true,
});
expect(expression(literal(size))).toMatchObject({
ok: false,
limited: true,
});
},
);

it('keeps malformed code invalid inside the larger body budget', () => {
const prefix = Array(300).fill('void 0;').join('\n');
expect(parseBody(`${prefix}\nreturn (;`)).toMatchObject({
ok: false,
message: expect.not.stringContaining('limit'),
});
});

it.each([MAX_PARSE_DEPTH - 1, MAX_PARSE_DEPTH])(
'accepts %i delimiters with a shallow AST',
(depth) => {
Expand Down Expand Up @@ -106,6 +129,41 @@ describe('parser analysis boundaries', () => {
});

describe('consumer parity at the admitted limits', () => {
it('admits forty bounded bodies without executing them and retains the node cap', async () => {
let executions = 0;
setCodeRunner(
createSandboxExecRunner(async () => {
executions++;
throw new Error('validation cannot execute a body');
}),
);
// 4096 tokens and 16384 code units per body, with a shallow, wide AST.
const body = `return !${array((MAX_BODY_PARSE_TOKENS - 4) / 2)};`;
const code = `/*${'x'.repeat(MAX_BODY_SOURCE_SIZE - body.length - 4)}*/${body}`;
const nodes = Array.from({ length: 40 }, (_, index) => ({
id: `step${index}`,
type: 'transform',
code,
}));
const document = {
version: 1,
name: 'bounded-bodies',
nodes,
output: nodes.map((node) => `{{ nodes.${node.id}.output }}`),
};
const admitted = await validate(document);
expect(admitted.errors).toEqual([]);
expect(admitted.warnings).toEqual([]);
const refused = await validate({
...document,
nodes: [...nodes, { id: 'step40', type: 'transform', code }],
});
expect(refused.errors).toEqual([
expect.objectContaining({ code: 'NODES_TOO_MANY' }),
]);
expect(executions).toBe(0);
});

it.each([MAX_SOURCE_SIZE - 2, MAX_SOURCE_SIZE - 1, MAX_SOURCE_SIZE])(
'keeps a quoted closer inside a %i-unit expression through runtime evaluation',
async (size) => {
Expand Down
19 changes: 15 additions & 4 deletions services/platform/lib/engine/core/syntax/parse.ts
Original file line number Diff line number Diff line change
Expand Up @@ -57,6 +57,10 @@ const OPTIONS: Options = {
export const MAX_SOURCE_SIZE = 8_192;
export const MAX_PARSE_DEPTH = 64;
export const MAX_PARSE_TOKENS = 512;
/** Transform bodies hold statements as well as expressions. Their separate
* bounded capacity does not enlarge templates, conditions or either depth cap. */
export const MAX_BODY_SOURCE_SIZE = 16_384;
export const MAX_BODY_PARSE_TOKENS = 4_096;
export const PARSE_LIMIT_MESSAGE =
'code exceeds the analysis size or depth limit';

Expand All @@ -69,14 +73,21 @@ function limit(start: number, end: number): ParseResult {
};
}

function budget(text: string, start: number, end: number): ParseResult | null {
if (end - start > MAX_SOURCE_SIZE) return limit(start, end);
function budget(
text: string,
start: number,
end: number,
kind: 'expression' | 'body' = 'expression',
): ParseResult | null {
const sourceLimit = kind === 'body' ? MAX_BODY_SOURCE_SIZE : MAX_SOURCE_SIZE;
const tokenLimit = kind === 'body' ? MAX_BODY_PARSE_TOKENS : MAX_PARSE_TOKENS;
if (end - start > sourceLimit) return limit(start, end);
let depth = 0;
let count = 0;
try {
for (const token of tokenizer(text.slice(start, end), OPTIONS)) {
const label = token.type.label;
if (++count > MAX_PARSE_TOKENS) return limit(start + token.start, end);
if (++count > tokenLimit) return limit(start + token.start, end);
if (['(', '[', '{', '${'].includes(label)) {
if (++depth > MAX_PARSE_DEPTH) return limit(start + token.start, end);
} else if ([')', ']', '}'].includes(label))
Expand Down Expand Up @@ -320,7 +331,7 @@ export function parseExpressionIn(
/** Parse a transform body the way the runner compiles it: a synchronous
* function body, so `return` is allowed at its top level. */
export function parseBody(code: string): ParseResult {
const limited = budget(code, 0, code.length);
const limited = budget(code, 0, code.length, 'body');
if (limited !== null) return limited;
try {
const ast = parse(code, { ...OPTIONS, allowReturnOutsideFunction: true });
Expand Down
Loading
Loading