Skip to content

fix: resolve markdown format loss in large docs - #623

Merged
pengfeixx merged 1 commit into
linuxdeepin:masterfrom
pengfeixx:agent/bug/fdc2336cdcf4
Sep 28, 2026
Merged

pengfeixx merged 1 commit into
linuxdeepin:masterfrom
pengfeixx:agent/bug/fdc2336cdcf4

Conversation

@pengfeixx

@pengfeixx pengfeixx commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

fix: resolve markdown format loss in large docs

  1. Root cause: progressive rendering split large documents into chunks
    and parsed each chunk independently, breaking cross-chunk markdown
    semantics like reference link definitions, loose lists, and tables
  2. Fix: parse entire document with parserCtx(md) first to preserve full
    markdown semantics, then insert parsed top-level nodes in batches via
    tr.insert with setTimeout between batches for UI responsiveness
  3. Impact: raise PROGRESSIVE_THRESHOLD from 64KB to 1MB so most documents
    use replaceAll fast path; only very large docs trigger progressive
    insertion with correct semantics

Log: Fix markdown format loss when pasting large markdown text in editor

Influence:

  1. Test pasting large markdown text (>1MB) with cross-chunk reference links
  2. Test pasting large markdown text with loose lists spanning chunks
  3. Test pasting large markdown text with tables split across chunks
  4. Verify normal-size markdown documents render correctly (fast path)
  5. Verify UI remains responsive during large document rendering

fix: 修复大段markdown文本渐进渲染格式丢失

  1. 根因:渐进渲染将大文档按顶层块切分后独立解析,导致跨块markdown
    语义断裂(引用式链接定义、松散列表、表格等被截断)
  2. 方案:先使用parserCtx(md)整篇解析保证语义完整,再将解析后的顶层
    节点分批通过tr.insert追加到文档末尾,每批间让出事件循环
  3. 影响:将渐进阈值从64KB提高至1MB,使绝大多数文档走replaceAll
    快速路径,仅超大文档触发渐进插入且语义正确

Log: 修复文本编辑器粘贴大段markdown文本时阅览区格式丢失问题

Influence:

  1. 测试粘贴超大markdown文本(>1MB)含跨块引用式链接定义
  2. 测试粘贴超大markdown文本含跨块松散列表
  3. 测试粘贴超大markdown文本含跨块表格
  4. 验证常规大小markdown文档渲染正确(快速路径)
  5. 验证大文档渲染期间界面保持响应

PMS: BUG-378329

Summary by Sourcery

Preserve Markdown formatting in large documents while keeping progressive rendering responsive.

Bug Fixes:

  • Preserve cross-block Markdown semantics when progressively rendering very large documents, preventing formatting loss in reference links, loose lists, tables, and similar constructs.

Enhancements:

  • Improve large-document rendering responsiveness by parsing the document as a whole and inserting parsed content in sized batches.
  • Raise the progressive-rendering threshold so documents up to 1 MB use the faster full replacement path while retaining prompt initial display for larger documents.

@sourcery-ai sourcery-ai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry @pengfeixx, you've used your own review budget of 250,000 diff characters for the last 7 days.

You can request another review in 1 day and 3 hours by commenting @sourcery-ai review. Upgrade to get a review now.

@sourcery-ai

sourcery-ai Bot commented Sep 28, 2026

Copy link
Copy Markdown

Reviewer's Guide

The PR fixes cross-chunk markdown formatting loss by parsing oversized documents as a whole, then progressively inserting the resulting top-level nodes in event-loop-yielding batches; the progressive path now activates only above 1MB while preserving initial rendering, progress reporting, scroll handling, and render-generation cancellation.

Sequence diagram for progressive large-document markdown rendering

sequenceDiagram
    participant Input as Markdown input
    participant Renderer as renderMarkdown
    participant Parser as parserCtx
    participant Editor as Editor state
    participant Timer as Event loop

    Input->>Renderer: renderMarkdown(md)
    alt md below 1MB
        Renderer->>Editor: replaceAll(md)
    else md at least 1MB
        Renderer->>Parser: parserCtx(md)
        Parser-->>Renderer: parsed top-level nodes
        Renderer->>Editor: replaceWithNodes(first batch)
        loop remaining batches
            Renderer->>Timer: setTimeout(step, 0)
            Timer->>Editor: appendChunk(batch)
        end
    end
Loading

Flow diagram for markdown rendering path selection

flowchart TD
    A[Markdown document] --> B{Size at least 1MB?}
    B -- No --> C[replaceAll fast path]
    B -- Yes --> D[Parse entire document with parserCtx]
    D --> E[Group parsed top-level nodes into batches]
    E --> F[Replace document with first batch]
    F --> G[Yield with setTimeout]
    G --> H[Insert next batch with tr.insert]
    H --> I{More batches?}
    I -- Yes --> G
    I -- No --> J[Complete rendering and restore scroll handling]
Loading

File-Level Changes

Change Details Files
Preserve markdown semantics by parsing the complete document before progressive rendering.
  • Remove heuristic source-text chunking and per-chunk parsing.
  • Parse the full document once through parserCtx(md) and retain top-level nodes for insertion.
src/editor/markdown/web/main.js
Implement responsive progressive insertion using batches of already-parsed nodes.
  • Batch nodes by estimated node size and replace the initial document with the first batch.
  • Append subsequent batches with transactions scheduled through setTimeout.
  • Track progress using parsed node sizes while retaining generation invalidation and scroll compensation.
src/editor/markdown/web/main.js
Reduce progressive-rendering frequency by raising the threshold to 1MB.
  • Increase PROGRESSIVE_THRESHOLD from 64KB to 1MB so normal-sized documents use the existing replaceAll fast path.
src/editor/markdown/web/main.js

Tips and commands

Interacting with Sourcery

  • Trigger a new review: Comment @sourcery-ai review on the pull request.
  • Continue discussions: Reply directly to Sourcery's review comments.
  • Generate a GitHub issue from a review comment: Ask Sourcery to create an
    issue from a review comment by replying to it. You can also reply to a
    review comment with @sourcery-ai issue to create an issue from it.
  • Generate a pull request title: Write @sourcery-ai anywhere in the pull
    request title to generate a title at any time. You can also comment
    @sourcery-ai title on the pull request to (re-)generate the title at any time.
  • Generate a pull request summary: Write @sourcery-ai summary anywhere in
    the pull request body to generate a PR summary at any time exactly where you
    want it. You can also comment @sourcery-ai summary on the pull request to
    (re-)generate the summary at any time.
  • Generate reviewer's guide: Comment @sourcery-ai guide on the pull
    request to (re-)generate the reviewer's guide at any time.
  • Resolve all Sourcery comments: Comment @sourcery-ai resolve on the
    pull request to resolve all Sourcery comments. Useful if you've already
    addressed all the comments and don't want to see them anymore.
  • Dismiss all Sourcery reviews: Comment @sourcery-ai dismiss on the pull
    request to dismiss all existing Sourcery reviews. Especially useful if you
    want to start fresh with a new review - don't forget to comment
    @sourcery-ai review to trigger a new review!

Customizing Your Experience

Access your dashboard to:

  • Enable or disable review features such as the Sourcery-generated pull request
    summary, the reviewer's guide, and others.
  • Change the review language.
  • Add, remove or edit custom review instructions.
  • Adjust other review settings.

Getting Help

@pengfeixx
pengfeixx force-pushed the agent/bug/fdc2336cdcf4 branch 2 times, most recently from b65d892 to fd919e4 Compare September 28, 2026 04:43
@deepin-ci-robot

Copy link
Copy Markdown

deepin pr auto review

🤖 AI 代码审查报告

总体评分: 76 分 (通过阈值: 70分)

Pass


📊 总体评价

项目 结果
审查结论 代码审查通过(有改进建议)
评分详情 代码正确修复了大文档Markdown渐进渲染格式丢失问题,无安全漏洞。但存在异常未捕获的逻辑缺陷和同步解析阻塞UI线程的性能问题,建议改进。

📋 PR 信息

项目 内容
PR #623
标题 fix: resolve markdown format loss in large docs
作者 pengfeixx
分支 agent/bug/fdc2336cdcf4 → master
提交 fd919e4
PMS BUG-378329
修改文件 2 个(分析 1 个,跳过 1 个编译产物)

🔍 详细分析

1. 语法逻辑 ❌

评价: 需改进 ❌ 不通过

潜在问题:

  1. src/editor/markdown/web/main.js:72-76 - renderProgressively() 函数中 ctx.get(parserCtx)(md) 调用缺少 try-catch 异常捕获。若解析器遇到异常输入抛出异常,错误将向上传播且未被处理。由于 lastValue 已在 renderMarkdown() 第54行更新为新内容,后续相同内容的调用会被第53行跳过,导致编辑器卡在旧内容无法恢复。

建议: 在 editor.action 回调中添加 try-catch,捕获异常后重置 lastValue 以允许重试:

let parsed = null;
editor.action((ctx) => {
    try {
        parsed = ctx.get(parserCtx)(md);
    } catch (e) {
        console.error('Markdown parse failed:', e);
        lastValue = '';
    }
});
if (!parsed) return;

2. 代码质量 ✅

评价: 良好 ✅ 通过

潜在问题:

  1. src/editor/markdown/web/main.js:96-114 - 当 nodes.length === 1 时,首批触发 replaceWithNodes + setTimeout(reapplyScroll),随后 index < nodes.length 为 false,else 分支又同步调用 reapplyScroll(),导致重复调用。虽然 reapplyScroll 是幂等的,但属于冗余。

建议: 注释详尽,结构清晰。建议优化单节点场景的冗余滚动调用,可在 else 分支前增加判断条件。


3. 代码性能 ❌

评价: 需改进 ❌ 不通过

潜在问题:

  1. src/editor/markdown/web/main.js:72-76 - 整篇文档通过 ctx.get(parserCtx)(md) 同步解析,对于超大文档(>1MB)会完全阻塞UI线程。代码注释中提到 5MB≈50s、10MB≈242s 的解析耗时,虽然阈值已提高至1MB使多数文档走快速路径,但超过1MB的文档在解析阶段仍会完全无响应。
  2. src/editor/markdown/web/main.js:78-82 - parsed.content.forEach 将所有顶层节点拷贝到 nodes 数组,创建了整个文档节点树的冗余内存副本。ProseMirror 的 Fragment 支持 child(i) 和 childCount 进行索引访问,可直接迭代而无需中间数组。

建议:

  1. 考虑使用 Web Worker 将初始解析移至后台线程,保持解析阶段UI响应
  2. 使用 parsed.content.child(i) 和 parsed.content.childCount 直接迭代,避免中间数组拷贝

4. 代码安全 🔒

评价: 优秀 ✅ 通过

🔐 发现 0 个安全漏洞

安全漏洞详情:
✅ 未发现安全漏洞

漏洞对比统计: 新增漏洞 0 个,减少漏洞 0 个,持平 0 个

建议: 无安全风险,代码安全合规。ProseMirror 解析器内置输入净化机制,无注入风险。


💡 改进建议代码示例

// 改进后的 renderProgressively 函数(添加异常处理 + 避免冗余拷贝)
function renderProgressively(md, gen) {
    let parsed = null;
    editor.action((ctx) => {
        try {
            parsed = ctx.get(parserCtx)(md);
        } catch (e) {
            console.error('Markdown parse failed:', e);
            lastValue = '';  // 重置以允许重试
            return;
        }
    });
    if (!parsed) return;

    // 直接使用 Fragment 索引访问,避免冗余数组拷贝
    const content = parsed.content;
    const totalSize = content.size;
    const nodeCount = content.childCount;
    let renderedSize = 0;
    let index = 0;
    buildProgress = { fraction: 0 };

    const step = () => {
        if (gen !== renderGeneration || !editor) return;
        if (index >= nodeCount) {
            buildProgress = null;
            if (!userScrolledDuringBuild) reapplyScroll();
            return;
        }
        const batch = content.child(index++);
        renderedSize += batch.nodeSize;
        if (index === 1) {
            replaceWithNodes([batch]);
            setTimeout(() => {
                if (gen !== renderGeneration) return;
                reapplyScroll();
            }, 0);
        } else {
            appendChunk([batch]);
        }
        buildProgress.fraction = totalSize > 0 ? renderedSize / totalSize : 1;
        if (index < nodeCount) {
            setTimeout(step, 0);
        } else {
            buildProgress = null;
            if (!userScrolledDuringBuild) reapplyScroll();
        }
    };
    step();
}

本报告由 AI 代码审查工具自动生成

@pengfeixx
pengfeixx force-pushed the agent/bug/fdc2336cdcf4 branch from fd919e4 to b6de264 Compare September 28, 2026 08:16
1. Root cause: progressive rendering split documents into chunks and
   parsed each independently, losing cross-chunk markdown semantics
   such as reference link definitions, tables, and loose lists
2. Fix: parse the entire document once with parserCtx(md) before batch
   inserting top-level nodes grouped by nodeSize (first batch ~256KB
   for fast first paint, subsequent batches ~1MB to bound dispatch
   count), ensuring all cross-chunk semantics are resolved correctly
   during progressive rendering
3. Impact: raised PROGRESSIVE_THRESHOLD from 64KB to 1MB so most
   documents use the fast replaceAll path; only >1MB documents trigger
   progressive insertion with correct whole-document parsing. Per-node
   batching filled a 1.4MB doc in ~172s; size-based batching fills it
   in ~7s (measured in headless Chromium)

Log: Fixed markdown format loss when pasting large text in editor

Influence:
1. Test pasting large markdown documents (>64KB) with tables and
   reference links to verify format renders correctly
2. Test normal-sized markdown documents for no regression
3. Test very large documents (>1MB) for progressive rendering behavior
   and verify the progressive fill completes within seconds

fix: 修复大段文本粘贴时阅览区 markdown 格式丢失

1. 根因:渐进渲染将文档分块后独立解析,导致跨块的 markdown 语义
   断裂,引用式链接定义、表格、松散列表等格式丢失
2. 方案:先用 parserCtx(md) 整篇解析文档,再将解析后的顶层节点
   按 nodeSize 分批(首批约 256KB 抢首屏、后续每批约 1MB 压低批数)
   插入,确保所有跨块语义正确解析
3. 影响:将渐进渲染阈值从 64KB 提高至 1MB,绝大多数文档走快速
   replaceAll 路径,仅超大文档触发渐进插入。逐节点分批填充 1.4MB
   文档实测约 172 秒,按大小分批后约 7 秒完成(headless Chromium 实测)

Log: 修复文本编辑器粘贴大段文本时预览区 markdown 格式丢失

Influence:
1. 测试粘贴含表格和引用链接的大段 markdown 文档(>64KB),验证格式正确渲染
2. 测试普通大小 markdown 文档无回归
3. 测试超大文档(>1MB)渐进渲染行为,验证渐进填充在秒级完成

PMS: BUG-378329
@pengfeixx
pengfeixx force-pushed the agent/bug/fdc2336cdcf4 branch from b6de264 to ead4518 Compare September 28, 2026 08:32
@deepin-ci-robot

Copy link
Copy Markdown

[APPROVALNOTIFIER] This PR is NOT APPROVED

This pull-request has been approved by: lzwind, pengfeixx

The full list of commands accepted by this bot can be found here.

Details Needs approval from an approver in each of these files:

Approvers can indicate their approval by writing /approve in a comment
Approvers can cancel approval by writing /approve cancel in a comment

@pengfeixx
pengfeixx merged commit 1551253 into linuxdeepin:master Sep 28, 2026
17 checks passed
@pengfeixx
pengfeixx deleted the agent/bug/fdc2336cdcf4 branch September 28, 2026 08:47
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants