Skip to content

feat: refine filename index data source validation - #266

Merged
Johnson-zs merged 1 commit into
linuxdeepin:develop/snipe-20260923from
Johnson-zs:develop/snipe-20260923
Sep 30, 2026
Merged

Johnson-zs merged 1 commit into
linuxdeepin:develop/snipe-20260923from
Johnson-zs:develop/snipe-20260923

Conversation

@Johnson-zs

@Johnson-zs Johnson-zs commented Sep 30, 2026 •

Copy link
Copy Markdown
  1. Enhanced the filename index validation for watch seeding and file enumeration operations
  2. Changed FSMonitorWorker's fast directory scan from using isFileNameIndexReadyForSearch() to new
    isFileNameIndexUsableAsDataSource() method
  3. Added disabled flag to persistent index state to track user's enable/ disable intent
  4. Updated index status store to handle disabled flag with proper persistence
  5. Modified FileNameIndexDBus::SetEnabled() to persistently record user's enable/disable decision
  6. Added comprehensive unit tests for new validation logic

The changes address a critical issue where the filename index validation was too strict for internal data source usage. Previously, the FSMonitorWorker used dfm-search's ready-for-search check which blocked fast directory scanning during normal index maintenance operations. The new validation distinguishes between search readiness (for external consumers) and data source usability (for internal operations), allowing index operations to continue during recovery/rebuild updates and backlog conditions while still blocking when the index is truly unusable.

Log: Refined filename index validation for better performance during maintenance operations

Influence:

  1. Test fast directory scan behavior when index is in various states (creating, updating, disabled, backlog exceeded)
  2. Verify user enable/disable functionality properly sets and persists disabled flag
  3. Test that directories created during update operations still get monitored correctly
  4. Verify that index continues to function during updateInProgress and backlogExceeded states
  5. Test index usability with corrupted or missing status.json files

feat: 优化文件名索引数据源验证逻辑

  1. 增强文件名索引验证机制,专门用于监视种子播撒和文件枚举操作
  2. 将FSMonitorWorker的快速目录扫描从使用isFileNameIndexReadyForSearch() 改为新的isFileNameIndexUsableAsDataSource()方法
  3. 向持久化索引状态添加禁用标志,以跟踪用户的启用/禁用意图
  4. 更新索引状态存储以正确处理禁用标志的持久化
  5. 修改FileNameIndexDBus::SetEnabled()来持久记录用户的启用/禁用决定
  6. 为新的验证逻辑添加全面的单元测试

这些变更解决了文件名索引验证对内部数据源使用过于严格的关键问题。之前,
FSMonitorWorker使用dfm-search的就绪搜索检查,在正常索引维护操作期间阻止
了快速目录扫描。新的验证区分了搜索就绪性(供外部消费者使用)和数据源可用
性(供内部操作使用),允许在恢复/重建更新和积压条件下继续索引操作,同时
在索引真正不可用时仍然阻止。

Log: 优化文件名索引验证,提升维护操作期间的性能表现

Influence:

  1. 测试索引处于各种状态(创建中、更新中、已禁用、积压超过限制)时的快速 目录扫描行为
  2. 验证用户启用/禁用功能是否正确设置和持久化禁用标志
  3. 测试在更新操作期间创建的目录是否仍能被正确监视
  4. 验证在updateInProgress和backlogExceeded状态下索引是否继续正常工作
  5. 测试索引在状态文件损坏或缺失时的可用性

Summary by Sourcery

Refine filename-index validation so internal data-source consumers can use partially stale indexes without unnecessary filesystem-wide fallback scans.

New Features:

  • Add a dedicated filename-index data-source usability check for watch seeding, enumeration, and related internal operations.
  • Persist the user's filename-index enabled or disabled intent in index status metadata.

Bug Fixes:

  • Allow fast directory scans to continue during recoverable update, backlog, and ordinary dirty states while still falling back when the index is unavailable, incomplete, incompatible, being created, or disabled.

Enhancements:

  • Update filename-index enablement handling and status persistence to preserve the disabled state without affecting other status fields.

Tests:

  • Add coverage for data-source usability across missing, invalid, disabled, creation, update, backlog, and dirty index states.
  • Verify disabled-state persistence, re-enablement, and compatibility with existing index status fields.

1. Enhanced the filename index validation for watch seeding and file
enumeration operations
2. Changed FSMonitorWorker's fast directory scan
from using isFileNameIndexReadyForSearch() to new
isFileNameIndexUsableAsDataSource() method
3. Added disabled flag to persistent index state to track user's enable/
disable intent
4. Updated index status store to handle disabled flag with proper
persistence
5. Modified FileNameIndexDBus::SetEnabled() to persistently record
user's enable/disable decision
6. Added comprehensive unit tests for new validation logic

The changes address a critical issue where the filename index validation
was too strict for internal data source usage. Previously, the
FSMonitorWorker used dfm-search's ready-for-search check which blocked
fast directory scanning during normal index maintenance operations.
The new validation distinguishes between search readiness (for external
consumers) and data source usability (for internal operations), allowing
index operations to continue during recovery/rebuild updates and backlog
conditions while still blocking when the index is truly unusable.

Log: Refined filename index validation for better performance during
maintenance operations

Influence:
1. Test fast directory scan behavior when index is in various states
(creating, updating, disabled, backlog exceeded)
2. Verify user enable/disable functionality properly sets and persists
disabled flag
3. Test that directories created during update operations still get
monitored correctly
4. Verify that index continues to function during updateInProgress and
backlogExceeded states
5. Test index usability with corrupted or missing status.json files

feat: 优化文件名索引数据源验证逻辑

1. 增强文件名索引验证机制,专门用于监视种子播撒和文件枚举操作
2. 将FSMonitorWorker的快速目录扫描从使用isFileNameIndexReadyForSearch()
改为新的isFileNameIndexUsableAsDataSource()方法
3. 向持久化索引状态添加禁用标志,以跟踪用户的启用/禁用意图
4. 更新索引状态存储以正确处理禁用标志的持久化
5. 修改FileNameIndexDBus::SetEnabled()来持久记录用户的启用/禁用决定
6. 为新的验证逻辑添加全面的单元测试

这些变更解决了文件名索引验证对内部数据源使用过于严格的关键问题。之前,
FSMonitorWorker使用dfm-search的就绪搜索检查,在正常索引维护操作期间阻止
了快速目录扫描。新的验证区分了搜索就绪性(供外部消费者使用)和数据源可用
性(供内部操作使用),允许在恢复/重建更新和积压条件下继续索引操作,同时
在索引真正不可用时仍然阻止。

Log: 优化文件名索引验证,提升维护操作期间的性能表现

Influence:
1. 测试索引处于各种状态(创建中、更新中、已禁用、积压超过限制)时的快速
目录扫描行为
2. 验证用户启用/禁用功能是否正确设置和持久化禁用标志
3. 测试在更新操作期间创建的目录是否仍能被正确监视
4. 验证在updateInProgress和backlogExceeded状态下索引是否继续正常工作
5. 测试索引在状态文件损坏或缺失时的可用性
@deepin-ci-robot

Copy link
Copy Markdown

[APPROVALNOTIFIER] This PR is NOT APPROVED

This pull-request has been approved by: Johnson-zs

The full list of commands accepted by this bot can be found here.

Details Needs approval from an approver in each of these files:

Approvers can indicate their approval by writing /approve in a comment
Approvers can cancel approval by writing /approve cancel in a comment

@sourcery-ai

sourcery-ai Bot commented Sep 30, 2026

Copy link
Copy Markdown

Reviewer's Guide

Refines filename-index validation by allowing internal watch seeding and enumeration to use substantially complete indexes during maintenance or backlog conditions, while blocking genuinely unusable or user-disabled indexes. Explicit user enable/disable intent is now persisted in status.json, with accompanying unit tests for validation and state persistence.

Sequence diagram for fast directory scan validation

sequenceDiagram
    participant W as FSMonitorWorker
    participant U as IndexUtility
    participant P as IndexProfile
    participant S as IndexStateStore
    W->>U: isFileNameIndexUsableAsDataSource()
    U->>P: filename()
    U->>P: isIndexAvailable()
    U->>S: isCompatibleVersion()
    U->>S: getLastUpdateTime()
    U->>S: isCreateInProgress()
    U->>S: isDisabled()
    alt usable data source
        U-->>W: true
        W->>W: Fast directory scan
    else unavailable, unbuilt, creating, or disabled
        U-->>W: false
        W->>W: Fallback traversal
    end
Loading

Sequence diagram for persistent filename index enablement

sequenceDiagram
    actor User
    participant D as FileNameIndexDBus
    participant S as IndexStateStore
    participant F as FSEventController
    User->>D: SetEnabled(enabled)
    D->>S: isDisabled()
    alt enablement intent changed
        D->>S: setDisabled(!enabled)
        S->>S: Persist disabled in status.json
    end
    D->>F: setEnabled(enabled)
Loading

Flow diagram for filename index data-source validation

flowchart TD
    A[Internal watch seeding or file enumeration] --> B["isFileNameIndexUsableAsDataSource()"]
    B --> C{Index available and compatible}
    C -->|No| D[Fallback traversal]
    C -->|Yes| E{Built and not creating}
    E -->|No| D
    E -->|Yes| F{User disabled}
    F -->|Yes| D
    F -->|No| G[Use filename index as data source]
    G --> H[Allow updateInProgress or backlogExceeded]
Loading

File-Level Changes

Change Details Files
Introduces a dedicated filename-index usability predicate for internal data-source operations, separate from external search readiness.
  • Allows fast directory scans and index-directory checks during update, backlog, and dirty states.
  • Rejects unavailable, incompatible, never-built, create-in-progress, or user-disabled indexes.
  • Adds unit coverage for status-file and transient-state behavior.
src/index/utils/indexutility.cpp
src/index/utils/indexutility.h
src/index/fsmonitor/fsmonitorworker.cpp
autotests/index/test_indexutility.cpp
autotests/index/test_filenameindexdbus.cpp
Persists the user’s filename-index enable/disable intent in index status metadata.
  • Adds a disabled status key with a false default for legacy or missing status files.
  • Uses read-modify-write persistence while preserving existing index fields.
  • Updates SetEnabled() to persist explicit user changes without marking normal shutdown as disabled.
  • Adds persistence and re-enable coverage.
src/index/index_global.h
src/index/state/indexstatestore.h
src/index/state/indexstatestore.cpp
src/index/dbus/filenameindexdbus.cpp
autotests/index/test_indexstatestore.cpp

Tips and commands

Interacting with Sourcery

  • Trigger a new review: Comment @sourcery-ai review on the pull request.
  • Continue discussions: Reply directly to Sourcery's review comments.
  • Generate a GitHub issue from a review comment: Ask Sourcery to create an
    issue from a review comment by replying to it. You can also reply to a
    review comment with @sourcery-ai issue to create an issue from it.
  • Generate a pull request title: Write @sourcery-ai anywhere in the pull
    request title to generate a title at any time. You can also comment
    @sourcery-ai title on the pull request to (re-)generate the title at any time.
  • Generate a pull request summary: Write @sourcery-ai summary anywhere in
    the pull request body to generate a PR summary at any time exactly where you
    want it. You can also comment @sourcery-ai summary on the pull request to
    (re-)generate the summary at any time.
  • Generate reviewer's guide: Comment @sourcery-ai guide on the pull
    request to (re-)generate the reviewer's guide at any time.
  • Resolve all Sourcery comments: Comment @sourcery-ai resolve on the
    pull request to resolve all Sourcery comments. Useful if you've already
    addressed all the comments and don't want to see them anymore.
  • Dismiss all Sourcery reviews: Comment @sourcery-ai dismiss on the pull
    request to dismiss all existing Sourcery reviews. Especially useful if you
    want to start fresh with a new review - don't forget to comment
    @sourcery-ai review to trigger a new review!

Customizing Your Experience

Access your dashboard to:

  • Enable or disable review features such as the Sourcery-generated pull request
    summary, the reviewer's guide, and others.
  • Change the review language.
  • Add, remove or edit custom review instructions.
  • Adjust other review settings.

Getting Help

@sourcery-ai sourcery-ai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hey - I've reviewed your changes and they look great!


Sourcery is free for open source - if you like our reviews please consider sharing them ✨

@deepin-ci-robot

Copy link
Copy Markdown

deepin pr auto review

AI 代码审查报告

PR: linuxdeepin/deepin-anything#266
标题: feat: refine filename index data source validation
作者: Johnson-zs
分支: develop/snipe-20260923 → develop/snipe-20260923
变更: 10 个文件,+236 行,-7 行
审查时间: 2026-09-30 15:25


总体评价

总分: 97/100 — 优秀

本次 PR 优化了文件名索引数据源验证逻辑,将 FSMonitorWorker 的快速目录扫描从使用 isFileNameIndexReadyForSearch() 改为新的 isFileNameIndexUsableAsDataSource() 方法,区分了搜索就绪性(外部消费者)和数据源可用性(内部操作)。同时新增了 disabled 标志的持久化机制,跟踪用户启停意图。代码质量高,注释详尽,测试覆盖全面,无安全漏洞。仅在 isFileNameIndexUsableAsDataSource() 函数中存在多次重复文件读取的性能优化空间。


四维度评分

维度 评分 状态 评价词
语法逻辑 25/25 ✓ 语法正确,逻辑清晰
代码质量 25/25 ✓ 代码结构清晰,注释完整
代码性能 17/20 ✓ 存在性能优化空间
代码安全 30/30 ✓ 存在0个安全漏洞

维度1:语法逻辑(25/25)✓

语法正确,逻辑清晰

变更分析

  1. isFileNameIndexUsableAsDataSource()(indexutility.cpp:29-41)

    • 新增函数,检查索引是否可作为数据源使用
    • 逻辑正确:依次检查索引可用性、版本兼容性、最后更新时间非空、非创建中、非禁用
    • 有意忽略 updateInProgress、backlogExceeded、dirty 状态,设计合理
  2. SetEnabled() 持久化逻辑(filenameindexdbus.cpp:279-290)

    • 条件 isDisabled() == enabled 正确识别状态变化时机:
      • enabled=true, isDisabled=true → 需要清除 disabled → setDisabled(false) ✓
      • enabled=true, isDisabled=false → 已启用,无需操作 ✓
      • enabled=false, isDisabled=false → 需要设置 disabled → setDisabled(true) ✓
      • enabled=false, isDisabled=true → 已禁用,无需操作 ✓
    • 仅在值变化时写盘,避免不必要的 I/O
  3. isDisabled()/setDisabled()(indexstatestore.cpp:200-211)

    • 采用与 isBacklogExceeded()/setBacklogExceeded() 一致的 read-modify-write 模式
    • 默认值 false(toBool(false)),向后兼容旧版文件

结论

无编译错误,无逻辑缺陷,边界处理完善。


维度2:代码质量(25/25)✓

代码结构清晰,注释完整

优点

  1. 注释质量优秀

    • indexutility.h 中 isFileNameIndexUsableAsDataSource() 的 Doxygen 注释详尽说明了设计意图:明确列出阻止条件和有意忽略的条件,解释了近似策略的权衡
    • indexstatestore.h 中 isDisabled()/setDisabled() 的注释说明了标志语义(用户意图 vs 任务状态)、行为细节和向后兼容性
    • filenameindexdbus.cpp 中 SetEnabled() 的中文注释解释了持久化机制和 cleanup 路径不误标的保证
    • fsmonitorworker.cpp 中的注释清晰阐述了从 isFileNameIndexReadyForSearch 到 isFileNameIndexUsableAsDataSource 的变更原因
  2. 代码结构合理

    • 函数职责单一,长度适中
    • isFileNameIndexUsableAsDataSource() 简洁明了,可读性好
    • 新增方法与现有 IndexStateStore 接口风格一致
  3. 测试覆盖全面

    • test_indexutility.cpp 新增 FileNameIndexUsableTest 测试套件,覆盖 9 种场景:
      • 缺失状态文件、就绪索引、空更新时间、版本不匹配、创建中、已禁用、更新中、积压超限、脏状态
    • test_indexstatestore.cpp 新增 Disabled_DefaultsFalseAndPersists 测试,验证默认值、持久化、清除及其他字段共存
    • test_filenameindexdbus.cpp 更新注释以反映新的验证逻辑
  4. 无残留调试代码

    • qWarning() 日志为合理的运行时警告,非调试残留

维度3:代码性能(17/20)✓

存在性能优化空间

性能问题

  1. isFileNameIndexUsableAsDataSource() 多次重复文件读取(indexutility.cpp:29-41)
    • 该函数调用 4 个 IndexStateStore 方法:isCompatibleVersion()、getLastUpdateTime()、isCreateInProgress()、isDisabled()
    • 每个方法独立调用 readStatusJson(statusFilePath()) 读取并解析 JSON 文件
    • 导致同一个小 JSON 文件被读取 4 次
    • 优化建议:读取一次 JSON 对象,在单个函数内完成所有检查:
      bool isFileNameIndexUsableAsDataSource()
      {
          const IndexProfile profile = IndexProfile::filename();
          if (!profile.isIndexAvailable())
              return false;
      
          const IndexStateStore store(profile);
          const QJsonObject obj = store.readStatusJson(store.statusFilePath());
          return store.isCompatibleVersion(obj)
                  && !obj.value(Defines::kLastUpdateTimeKey).toString().isEmpty()
                  && !obj.value(Defines::kCreateInProgressKey).toBool(false)
                  && !obj.value(Defines::kDisabledKey).toBool(false);
      }
    • 影响评估:该函数在启动路径(watch 播种)和路径检查(isIndexWithAnything)中调用,非高频热路径,影响有限

说明

  • 该模式与 IndexStateStore 现有接口设计一致(所有 getter 方法均独立读取文件)
  • 优化需要修改 IndexStateStore 的公共接口或在调用方内联 JSON 解析,属于架构层面改进
  • SetEnabled() 中 isDisabled() + setDisabled() 产生 2 次读取,但为用户触发操作,影响可忽略

维度4:代码安全(30/30)✓

存在0个安全漏洞

安全审查

  1. 文件路径处理:statusFilePath() 返回固定路径,无用户可控输入,无路径遍历风险
  2. JSON 处理:使用 Qt 标准 JSON 类(QJsonObject/QJsonDocument),无反序列化漏洞
  3. DBus 接口:SetEnabled(bool) 参数为简单布尔类型,无注入风险;该方法为已有 DBus 接口,本次仅新增 disabled 标志持久化
  4. 敏感信息:无硬编码密钥、密码或 Token
  5. 命令注入:无系统调用或命令执行
  6. 并发安全:read-modify-write 模式与现有代码一致,无新增竞态风险

扫描器误报说明

安全扫描器在 index_global.h 第 57-58 行报告 2 个"硬编码密钥"漏洞,经人工审查确认为误报:

  • 第 57 行:kLightIncrementFileCountThreshold = QLatin1String("lightIncrementFileCountThreshold") — JSON 键名常量
  • 第 58 行:kLightIncrementOcrFileCountThreshold = QLatin1String("lightIncrementOcrFileCountThreshold") — JSON 键名常量
  • 这两行均为已有代码,不在本次 diff 范围内
  • 本次 diff 在该文件仅新增 kDisabledKey = QLatin1String("disabled"),同样是 JSON 键名常量

漏洞对比统计:新增漏洞 0 个,减少漏洞 0 个,持平 0 个


改进建议

1. 性能优化:减少重复文件读取(轻微)

文件: src/index/utils/indexutility.cpp 第 29-41 行
函数: isFileNameIndexUsableAsDataSource()

当前实现通过 4 个 IndexStateStore 方法各自读取 JSON 文件,可考虑读取一次并在调用方完成检查,或为 IndexStateStore 增加批量查询接口。

2. 可读性优化:SetEnabled 条件表达式(轻微)

文件: src/index/dbus/filenameindexdbus.cpp 第 285 行
函数: SetEnabled()

当前条件 isDisabled() == enabled 逻辑正确但语义不够直观,可考虑使用更具描述性的写法:

const bool desiredDisabled = !enabled;
if (d->runtime->stateStore().isDisabled() != desiredDisabled) {
    d->runtime->stateStore().setDisabled(desiredDisabled);
}

审查结论

本次 PR 是一次高质量的功能优化,核心变更是将文件名索引验证从"搜索就绪性"细化为"数据源可用性",允许索引在恢复/重建更新期间继续支持快速目录扫描。新增的 disabled 标志持久化机制设计合理,与现有 IndexStateStore 架构一致。测试覆盖全面,注释详尽。仅存在轻微的性能优化空间(重复文件读取),不影响功能正确性。

建议: 可以合并,建议考虑上述性能优化建议。

@Johnson-zs
Johnson-zs merged commit 523b104 into linuxdeepin:develop/snipe-20260923 Sep 30, 2026
17 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants