Skip to content

feat: add skip_store setting - #1773

Open
Nexisato wants to merge 26 commits into
mainfrom
feat/cloud-only
Open

Nexisato wants to merge 26 commits into
mainfrom
feat/cloud-only

Conversation

@Nexisato

@Nexisato Nexisato commented Sep 8, 2026 •

Copy link
Copy Markdown
Collaborator

新增 core.skip_store 设置(仅 online 模式合法):开启后 SDK 不在本地产生任何文件(无 swanlog/、run-*.swanlab、media/、files/、debug/),全部数据直传云端。

Related Issue: #1713

使用

swanlab.init(settings=swanlab.Settings(core=swanlab.Settings.Core(skip_store=True)))
# 或环境变量 SWANLAB_CORE_SKIP_STORE=true

设计

协议:proto 全部追加字段,向后兼容——MediaItem.payload=5(optional,区分空文件与缺失)、SaveRecord.payload=6、CoreSettings.skip_store=10、ProbeSettings.skip_store=13。本地 DataStore 文件格式版本未变,旧 swanlog 仍可被新版读取与 sync。

落盘短路:DataStoreWriter(skip=True) 的 open/write/close 全部无副作用(write 仅计数供 close 统计);init 跳过一切目录创建;诊断日志只输出终端。

媒体:transform(path=None) 时将内容写入 MediaItem.payload 而非落盘;sender 经 MemoryViewReader 零拷贝内存直传对象存储,不触碰本地 media 路径。

内部 save(config/metadata/requirements/conda):内容按落盘同款编码内联进 SaveRecord.payload(config 复用 dump_config,避免云端结构漂移);sender 解析后走 profile 上传。解析失败属确定性脏数据,告警跳过,不进入 Transport 无限重试;上传失败保持既有 ApiError 分类(5xx 重试 / 4xx 跳过)。

CUSTOM save:payload 恒空(非空视为协议违约丢弃),从 source_path 原路径读取上传,不创建本地软链接镜像;policy="live" 的 watcher 直接监听源文件所在目录(direct-source 模式,事件路径精确匹配,同源多 name 一对多注册)。

probe:通过新增的显式 ProbeSettings.skip_store 字段感知模式,metadata/requirements/conda 直接注入 payload,不写 files/ 目录。

模式约束:Settings validator 保证 skip_store 仅 online 合法(含 cloud 别名归一化、env 注入、merge 降级路径);交互式引导从 online 降级到 offline 时显式关闭 skip_store 并告警,保证离线数据正常落盘。

取舍

  • 无本地副本:进程崩溃后未确认的 record 不可恢复,swanlab sync / swanlab watch 不适用;init 时打印警告明示。
  • 媒体内容在上传确认前驻留内存(RecordBuffer 积压时随之增长)。

测试

单元测试与静态检查

项目 命令 结果
全量单测 uv run pytest tests/unit 1861 passed, 21 skipped
Lint uv run ruff check . All checks passed
类型检查 uv run basedpyright 0 errors, 0 warnings
Go 生成代码 cd core && go build ./... OK
proto 生成 make proto Python/Go 与 protos/ 源同步,无额外 diff

覆盖点:

  • e2e:全程零本地文件 + 标量/日志/媒体/三类 save 全量上云断言;交互式降级关闭 skip_store
  • 单测:store skip 计数、sender payload 双源与脏数据跳过、watcher direct-source、config 事件内联、settings 校验矩阵

基准测试

新增两个基准(tests/benchmark/),分别衡量本地持久化层与端到端运行时链路的开销差异。

1. 本地持久化层

tests/benchmark/sdk/internal/core_python/store/bench_store_skip.py,直接走生产路径 CorePython._store_records,对比落盘(skip_store=False)与完全跳过持久化(skip_store=True)的本地工作:

  • 标量:200 key × 5000 step = 1,000,000 条(与 bench_metrics_steps.py 对齐),落盘写 LevelDB log,skip 仅计数
  • 媒体:1,000 条 × 32 KiB(对应 Image.transform),落盘写 media/image/(含 fsync),skip 内联 MediaItem.payload
  • 文件保存:100 个 CUSTOM save(对应 Core._handle_custom_save),落盘建软链接镜像,skip 不建
uv run pytest tests/benchmark/sdk/internal/core_python/store/bench_store_skip.py -v -s

本机(macOS / Apple Silicon,Python 3.11,每类重复 3 次取最优):

类型 规模 落盘 跳过 加速比
标量 200 key × 5000 step 0.5071 s 0.003420 s 148.3x
媒体 1,000 × 32 KiB 0.2219 s 0.000009 s ~24,430x
文件保存 100 个 CUSTOM save 0.0176 s 0.000043 s 409.7x
合计 — 0.7466 s 0.003473 s 215.0x

落盘侧附加指标:标量吞吐 1.97 M rec/s、媒体写入带宽 140.8 MiB/s、run-*.swanlab 34.4 MB、媒体文件 1,000 个 / 32,768,000 B、save 镜像 100 个;skip 侧无任何本地文件。数据完整性:标量可完整回读,媒体文件数/字节数、save 镜像数均与写入一致。

2. 端到端(online + mock HTTP)

tests/benchmark/sdk/cmd/bench_skip_store_e2e.py,完整跑 swanlab.init → log/log_image/save → finish,HTTP 全部 mock,对比 skip_store 对用户线程延时与端到端吞吐量的影响。为让 producer / finish 两阶段边界确定,record_interval 设为大值,上传统一发生在 finish。

负载:1000 step × 100 key = 100,000 标量、250 张媒体、50 个 save。

uv run pytest tests/benchmark/sdk/cmd/bench_skip_store_e2e.py -v -s

本机(同上,每场景重复 2 次取最优):

指标 落盘 skip_store 变化
端到端总耗时 1.5225 s 1.3576 s -10.8%
finish 排空 + 上传 1304 ms 517 ms -60.4%
端到端吞吐量 65,683 rec/s 73,657 rec/s +12.1%
producer 阶段墙钟 0.195 s 0.835 s +328%
run.log 平均延时 113.4 µs 117.8 µs ~持平
run.log p50 / p95 / p99 113.4 / 129.6 / 171.9 µs 116.3 / 135.8 / 160.0 µs 持平
本地产物 6.18 MB / 250 media / 50 links 0 —

解读:

  • 端到端收益主要来自 finish 排空与上传阶段(本地 store/媒体落盘的 I/O 被消除),整体墙钟下降约 11%。
  • 用户线程单次 run.log 延时基本不变(~113 µs),因为生产端本就只做入队。
  • producer 阶段墙钟反而变长,是 GIL 竞争的副作用:skip 下 BackgroundConsumer 的本地 store 从磁盘 I/O(让出 GIL)变成纯内存 Python 操作(占用 GIL),与主线程竞争更激烈。总工作量仍是下降的,故总耗时与吞吐量均改善。
  • 这里 HTTP 已 mock,上传是本机内存操作;真实网络/磁盘环境下 finish 的绝对耗时会更大,skip 的 I/O 节省占比通常更可观。

注:本地持久化层基准衡量的是「protobuf 序列化 + 缓冲写」这一跳;媒体落盘经 fs.safe_write 含 fsync,媒体行的加速比同时包含跳过 fsync 的收益。绝对数值随机器变化,加速比与吞吐量级为主要参考。

@Nexisato Nexisato self-assigned this Sep 10, 2026
@Nexisato
Nexisato requested a review from SAKURA-CAT September 11, 2026 05:22
@Nexisato Nexisato added the 💪 enhancement New feature or request label Sep 14, 2026

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

这个sender有点重量级了,一些函数可以拆分一下,可以重构为:
transport下有个sender文件夹,然后导出HttpRecordSender,这样测试也好写一些,不过当前PR的目的不是这个,所以可以先写个issue记录,等这个PR合并后再拆分

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

对DataStore的改动按我理解是为了避免调用方加一堆If else

不过仓库内既有风格是 Null Object 模式(swanlab/sdk/internal/run/components/null.py 的 NullEmitter/NullConsumer/NullTerminalProxy),writer 的 skip_store 应改用同样方式——新建 NullDataStoreWriter(no-op open/write/skip_records/close),由 core.py 按 skip_store 选择实例,去掉 DataStoreWriter 内部的 skip 标志、write() assert 和 skip_records() 条件分支。

Null Object 模式的好处为避免业务类型内部处理一些非必要 if else,职责也更明确

"""direct-source 模式:不依赖本地镜像目录,直接监听 source_path 所在目录。

- 按 source_path.parent 分组,一个目录只 schedule 一次;
- _registered 以源文件绝对路径为 key,事件精确匹配(同目录其他文件被忽略);

@SAKURA-CAT SAKURA-CAT Oct 5, 2026 •

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

这里可以再评估一下为了skip store模式单独开发 direct-source 模式的合理性
主要是会不会引来一些不必要的bug,如果这里不太确定,我的建议是不用添加watcher,和对writer的处理一样新建一个 Null Object
然后在save层warning,只允许now,通过限制行为的方式来减少工作量,并且减少不确定性

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

同上一个评论,如果这里很复杂的话也没必要加一些魔法 hhh

raise TypeError("Object has no len")


class MemoryViewReader(io.RawIOBase):

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

这个对象的意义是啥

source_path=ctx.metadata_file.absolute().as_posix(),
type=SaveType.SAVE_TYPE_METADATA,
)
content = sys_info.metadata.model_dump_json(by_alias=True)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

probe里面对metadata、requirements、conda有重复的处理逻辑,可以写个helper函数(例如 make_save_record)来解决,这样也能写测试

至于config、builder那边,由于跨模块了,另行处理

# 也保证 Windows 下文件内容 (LF) 与上方 sha256/size 计算结果一致
fs.safe_write(path / filename, content_encode, mode="wb")
return MediaItem(filename=filename, sha256=sha256, size=len(content_encode), caption=self.caption)
item = MediaItem(

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

笑死了 感觉可以封装个函数

在 TransformMedia 基类(swanlab/sdk/internal/run/transforms/init.py 或所在基类模块)加一个终态助手:

def _attach_content(self, item: MediaItem, path: Optional[Path], content: bytes) -> MediaItem:
    """skip_store(path=None)时内容进 payload(空内容也赋值以保留 presence),否则落盘。"""
    if path is None:
        item.payload = content
    else:
        fs.safe_write(path / item.filename, content, mode="wb")
    return item

return self

@model_validator(mode="after")
def validate_skip_store(self) -> "Settings":

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

这个加上了,init那边的校验还有意义嘛,看起来两者冲突

Extract a `should_mkdirs` flag to avoid duplicating the mode/skip_store
check, and clarify the affected comments and Go formatting in the
generated save proto.
@SAKURA-CAT

SAKURA-CAT commented Oct 5, 2026 •

Copy link
Copy Markdown
Member

顺便,测试似乎失败了,看了一下似乎又是windows上的特殊行为,可以记个issue后续修复:

根因分析

失败测试:test_watch_sources_missing_source_keeps_registration (tests/unit/.../test_watcher.py:145)

compute_signature(swanlab/sdk/internal/core_python/watcher/helper.py:35)的签名是 mtime_ns + size:

st = os.stat(path)
return f"{st.st_mtime_ns}:{st.st_size}"

测试流程是:写 b"v1" → unlink → 写 b"v2" → _process_change。

问题在于:

  1. b"v1" 和 b"v2" size 相同(2 字节)
  2. Windows 上删除后立刻重建文件,新文件拿到的时间戳落在 NTFS 时间戳粒度内(同 tick 内 mtime_ns 可能完全相同),且测试文件注释里已记录过 Windows 的 stat 缓存问题(test_watcher.py:32-34 的 REAL_OBSERVER_DEBOUNCE 注释:NTFS 目录项缓存约 1s 内读到旧 stat)

结果 new_sig == entry.signature,_process_change (watcher/init.py:153) 判定"签名未变"而跳过回调 → assert_called_once 报 called 0 times。Linux/macOS 上 mtime 分辨率高,不会复现。

修复方案(计划)

推荐:签名中加入文件标识 st_ino(Windows 上 NTFS 也提供有效 file index),使"删除后重建"即使 mtime/size 都相同也必然产生不同签名:

# helper.py compute_signature
st = os.stat(path)
return f"{st.st_mtime_ns}:{st.st_size}:{st.st_ino}"
  • 覆盖测试场景(重建后 inode 变化)且语义正确:删除重建本质上是新文件
  • 不影响真实 observer 的两个跨平台测试(它们依赖回调触发,签名变严只会更早判定"变化",不会漏报)
  • 代码库中 compute_signature 只有 watch/watch_sources/_process_change 三个调用方,都在 watcher 内部,无外部影响

备选(不推荐):只改测试用不同长度的内容(如 b"v2-longer"),能绕过本例但未修复真实场景——生产中用户以相同 size 覆盖保存(如定点数 checkpoint)在 Windows 上同样会漏报。

验证:uv run pytest tests/unit/sdk/internal/core_python/watcher/(本地 macOS 无法复现 Windows 失败,需推送后看 CI Windows 矩阵)。

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

💪 enhancement New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants