Skip to content

[Data Importer]: Interpretation-stage events (PreInterpretFileEvent, PreQueueRowEvent) to modify/skip/fan-out rows without custom interpreters #340

Description

@alexbaat

Affected capability

DataHub (Data Importer)

Feature description

Proposal: dispatch events at the interpretation stage of the Data Importer (file → rows), so projects can adjust rows without writing a custom interpreter.

Problem

Today the whole path from source file to queue item runs inside AbstractInterpreter with no extension point:

  • interpretFile() → doInterpretFileAndCallProcessRow() → processImportRow() → QueueService::addItemToQueue()

The only way to influence this stage is a custom interpreter. All shipped interpreters are final (and AbstractInterpreter is @internal), so even a small row-level adjustment means copying the whole CSV/XLSX/JSON/XML reading logic plus fileValid(), previewData(), setSettings(), plus a Studio UI dynamic type so the type is selectable — hundreds of lines of boilerplate for a few lines of actual logic.

The mapping/transformation stage cannot cover these cases:

  • Skipping a row — impossible: by the time operators run, Resolver::loadOrCreateAndPrepareElement() has already created/loaded the element; operators have no veto channel (throwing produces an error log entry per row and fires ProcessElementExceptionEvent).
  • Fan-out (1 source row → N elements) — impossible by construction: one queue item is one processElement() is one save().
  • Cross-row state (e.g. SAP report exports where a group header row applies to all following rows) — impossible: rows are isolated queue items, possibly processed in parallel.
  • Synthetic columns for the resolver (computed key/path columns) — must exist in the raw row before the resolver's identifier extraction and the path/location strategies run, i.e. before mapping.

Real-world evidence: our project ships three custom interpreters (~840 lines PHP + a Studio UI module-federation remote) whose entire raison d'être is: filter rows by a value blacklist, carry a group marker forward across rows + add synthetic key/path columns, and pivot one wide row into up to three normalized rows. Each copies the CSV interpreter because it is final.

Proposed solution

Two new events (PR follows in pimcore/data-importer):

  1. PreQueueRowEvent — dispatched in AbstractInterpreter::processImportRow() for every extracted row, before the delta check and queueing. Listeners can:

    • modify the row: $event->setRows([$changedRow])
    • skip it: $event->skipRow() / $event->setRows([])
    • fan it out: $event->setRows([$rowA, $rowB, $rowC]) (each queued and imported as its own element)
    • skipRow(keepInCleanupIdentifierCache: true) skips the row but still registers its identifier, so an active cleanup strategy does not delete/unpublish the row's existing element.
  2. PreInterpretFileEvent — dispatched at the start of interpretFile() with a settable path, so listeners can normalize the file (transcode, strip a report preamble, rewrite delimiters) and stateful row listeners get a per-run reset signal.

Both events are also dispatched (flagged with isPreview()) in the Studio preview and column-header code paths, so the mapping UI shows exactly the columns an actual import produces — including listener-added synthetic columns.

Wiring: the interpreter compiler pass adds a setEventDispatcher() call to every tagged interpreter service, so built-in and custom interpreters dispatch the events; no constructor BC break (AbstractInterpreter's constructor is unchanged), and without listeners behavior is byte-for-byte identical.

This follows the invitation in the Data Importer events docs ("More events to come when needed (just provide PRs ;-)") and complements pimcore/data-importer#607, which adds a similar skip capability at the save stage.

Alternative(s)

  • Keep writing custom interpreters (status quo): large copied boilerplate per project, breaks on upstream changes to the copied code, needs a Studio UI remote per interpreter.
  • Un-final the shipped interpreters + make AbstractInterpreter public API: larger BC surface than two events.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Fields

    Affected capability

    None yet

    Galaxy

    None yet

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions