Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -124,7 +124,7 @@ Fully modular pipeline from document parsing, vectorization, and retrieval to LL
| Knowledge Base Types | FAQ / Document / Wiki with folder import, URL import, multi-tag management, and online entry |
| Per-Upload Process Config | Override parser, chunking, multimodal (VLM / ASR), graph extraction, and question generation per upload batch via upload-confirm dialog or `process_config` API; reparse with new settings |
| Batch Reparse | Re-queue parsing for multiple documents at once with optional per-batch `process_config` |
| Data Source Import | Auto-sync from Feishu / Notion / Yuque / RSS feeds (more data sources coming soon); incremental and full sync |
| Data Source Import | Auto-sync from Feishu / DingTalk / Notion / Yuque / RSS feeds (more data sources coming soon); incremental and full sync |
| Document Formats | PDF / Word / Txt / Markdown / HTML / EPUB / MHTML / Images / CSV / Excel / PPT / JSON |
| Retrieval Strategies | BM25 sparse / Dense retrieval / GraphRAG / parent-child chunking / HNSW-accelerated pgvector (1024-dim) / multi-dimensional indexing |
| Batch Selection | Marquee drag-select multiple documents in the KB list for batch operations |
Expand Down
2 changes: 1 addition & 1 deletion README_CN.md
Original file line number Diff line number Diff line change
Expand Up @@ -123,7 +123,7 @@
| 知识库类型 | FAQ / 文档 / Wiki,支持文件夹导入、URL 导入、多标签管理、在线录入 |
| 按批次解析配置 | 上传确认对话框或 `process_config` API 覆盖解析引擎、分块、多模态(VLM / ASR)、图谱抽取与问题生成;支持 reparse 时调整配置 |
| 批量重新解析 | 一次为多篇文档重新排队解析,可携带批次级 `process_config` |
| 数据源导入 | 飞书 / Notion / 语雀 / RSS 订阅自动同步(更多数据源开发中),支持增量与全量同步 |
| 数据源导入 | 飞书 / 钉钉 / Notion / 语雀 / RSS 订阅自动同步(更多数据源开发中),支持增量与全量同步 |
| 文档格式 | PDF / Word / Txt / Markdown / HTML / EPUB / MHTML / 图片 / CSV / Excel / PPT / JSON |
| 检索策略 | BM25 稀疏召回 / Dense 稠密召回 / GraphRAG 图谱增强 / 父子分块 / pgvector HNSW 加速(1024 维)/ 多维度索引 |
| 批量选择 | 知识库文档列表支持框选(marquee)多选,便于批量操作 |
Expand Down
2 changes: 1 addition & 1 deletion README_JA.md
Original file line number Diff line number Diff line change
Expand Up @@ -124,7 +124,7 @@ Feishu、Notion、Yuqueなどの外部プラットフォームからのナレッ
| ナレッジベースタイプ | FAQ / ドキュメント / Wiki、フォルダーインポート・URL インポート・複数タグ管理・オンライン入力 |
| アップロード単位の解析設定 | アップロード確認ダイアログまたは `process_config` API でパーサー・チャンキング・マルチモーダル(VLM / ASR)・グラフ抽出・質問生成をバッチ単位で上書き;reparse 時も設定変更可能 |
| 一括 reparse | 複数ドキュメントの解析を一度に再キュー、バッチ単位の `process_config` 対応 |
| データソースインポート | Feishu / Notion / Yuque / RSS フィードの自動同期(他のデータソースも開発中)、増分・全量同期対応 |
| データソースインポート | Feishu / DingTalk / Notion / Yuque / RSS フィードの自動同期(他のデータソースも開発中)、増分・全量同期対応 |
| 文書フォーマット | PDF / Word / Txt / Markdown / HTML / EPUB / MHTML / 画像 / CSV / Excel / PPT / JSON |
| 検索戦略 | BM25 疎検索 / Dense 密検索 / GraphRAG グラフ強化 / 親子チャンキング / pgvector HNSW 加速(1024 次元)/ 多次元インデックス |
| 一括選択 | KB リストでマーキー(ドラッグ)複数選択によるバッチ操作 |
Expand Down
2 changes: 1 addition & 1 deletion README_KO.md
Original file line number Diff line number Diff line change
Expand Up @@ -133,7 +133,7 @@ Feishu, Notion, Yuque 등 외부 플랫폼에서 지식 자동 동기화를 지
| 지식베이스 타입 | FAQ / 문서 / Wiki, 폴더 임포트·URL 임포트·다중 태그 관리·온라인 입력 |
| 업로드 단위 파싱 설정 | 업로드 확인 대화상자 또는 `process_config` API로 파서·청킹·멀티모달(VLM / ASR)·그래프 추출·질문 생성을 배치 단위로 덮어쓰기; reparse 시 설정 변경 지원 |
| 일괄 reparse | 여러 문서의 파싱을 한 번에 재큐잉, 배치 단위 `process_config` 지원 |
| 데이터 소스 임포트 | Feishu / Notion / Yuque / RSS 피드 자동 동기화(추가 데이터 소스 개발 중), 증분·전체 동기화 지원 |
| 데이터 소스 임포트 | Feishu / DingTalk / Notion / Yuque / RSS 피드 자동 동기화(추가 데이터 소스 개발 중), 증분·전체 동기화 지원 |
| 문서 포맷 | PDF / Word / Txt / Markdown / HTML / EPUB / MHTML / 이미지 / CSV / Excel / PPT / JSON |
| 검색 전략 | BM25 희소 / Dense 밀집 / GraphRAG 그래프 강화 / 부모-자식 청킹 / pgvector HNSW 가속(1024차원) / 다차원 인덱싱 |
| 일괄 선택 | KB 목록에서 마키(드래그) 다중 선택으로 일괄 작업 |
Expand Down
16 changes: 14 additions & 2 deletions docs/wiki/集成扩展/数据源导入开发.md
Original file line number Diff line number Diff line change
@@ -1,13 +1,13 @@
---
title: 数据源导入开发
tags: [集成扩展, 数据源, 飞书, 同步, 连接器]
tags: [集成扩展, 数据源, 飞书, 钉钉, 同步, 连接器]
aliases: [数据源导入, DataSource, 数据同步]
source: 数据源导入开发文档.md
---

# 数据源导入开发

WeKnora 的数据源导入模块支持从外部平台(飞书、企业微信、Notion、Confluence 等)自动导入和同步内容到知识库。用户可配置数据源连接,选择需要同步的资源,并通过手动触发或定时调度自动完成内容的增量/全量同步。
WeKnora 的数据源导入模块支持从外部平台(飞书、钉钉、Notion、语雀、RSS 等)自动导入和同步内容到知识库。用户可配置数据源连接,选择需要同步的资源,并通过手动触发或定时调度自动完成内容的增量/全量同步。

数据源绑定到知识库,一个知识库可接入多个数据源。凭证使用 AES-256-GCM 加密存储。

Expand All @@ -18,6 +18,7 @@ WeKnora 的数据源导入模块支持从外部平台(飞书、企业微信、
| 连接器 | 认证方式 | 增量同步 | 删除同步 |
|--------|---------|:-:|:-:|
| 飞书 (Feishu) | OAuth2 (Tenant Access Token) | ✅ | ✅ |
| 钉钉文档 (DingTalk Docs) | OAuth2 App Access Token | ✅ | ✅ |

## 快速接入:飞书知识库

Expand All @@ -29,6 +30,17 @@ WeKnora 的数据源导入模块支持从外部平台(飞书、企业微信、

> 注意:飞书国际版(Lark)同样支持,自动适配 `https://open.larksuite.com` 的 API 地址

## 快速接入:钉钉文档

1. 创建钉钉企业内部应用并获取 Client ID / Client Secret
2. 开通知识库读、知识库节点读、企业存储文件读权限并发布应用版本
3. 准备一个能够读取目标知识库的用户 UnionID
4. 在知识库设置页添加“钉钉文档”,测试连接并选择空间、文件夹或单篇文档
5. 配置全量或增量同步并触发首次同步

连接器使用官方知识库节点和文档块 API,当前仅同步 `FILE + ALIDOC + adoc`
在线文档。目录遍历不完整时不会判定删除,以避免临时权限或网络异常造成误删除。

## 架构设计

```
Expand Down
36 changes: 35 additions & 1 deletion docs/数据源导入开发文档.md
Original file line number Diff line number Diff line change
@@ -1,13 +1,14 @@
# 数据源导入开发文档

WeKnora 的数据源导入模块支持从外部平台(飞书、企业微信、Notion、Confluence 等)自动导入和同步内容到知识库。用户可配置数据源连接,选择需要同步的资源,并通过手动触发或定时调度自动完成内容的增量/全量同步。
WeKnora 的数据源导入模块支持从外部平台(飞书、钉钉、Notion、语雀、RSS 等)自动导入和同步内容到知识库。用户可配置数据源连接,选择需要同步的资源,并通过手动触发或定时调度自动完成内容的增量/全量同步。

数据源绑定到知识库,一个知识库可接入多个数据源。所有配置通过前端知识库设置页面管理,凭证使用 AES-256-GCM 加密存储在数据库中。

## 目录

- [快速接入指南](#快速接入指南)
- [飞书知识库接入](#飞书知识库接入)
- [钉钉文档接入](#钉钉文档接入)
- [前端管理](#前端管理)
- [架构总览](#架构总览)
- [数据模型](#数据模型)
Expand Down Expand Up @@ -128,6 +129,39 @@ Lark 是飞书的国际版。Wiki / docx / drive 接口与飞书一致,**共
若你此前用「飞书」连接器加 `base_url=https://open.larksuite.com` 的方式接入过 Lark,该配置
仍然有效(`base_url` 保留为显式覆盖项),但新建数据源请直接选「Lark」类型。

### 钉钉文档接入

钉钉连接器通过官方服务端 API 同步知识库中的在线文档。当前支持选择整个知识库、
文件夹子树或单篇在线文档,支持全量和基于节点 `modifiedTime` 的增量同步。

#### 第一步:创建企业内部应用

1. 登录 [钉钉开放平台](https://open-dev.dingtalk.com/) 并创建企业内部应用
2. 获取 **Client ID(AppKey)** 与 **Client Secret(AppSecret)**
3. 创建并发布应用版本

#### 第二步:开通只读权限

按照钉钉官方 API 的权限要求开通:

- 知识库读权限
- 知识库节点读权限
- 企业存储文件读权限

#### 第三步:准备操作人 UnionID

知识库与文档接口要求传入 `operatorId`。该 UnionID 对应的用户必须能够读取目标知识库;
应用权限和操作人资源权限缺一不可。

#### 第四步:添加数据源

进入 **知识库设置 → 数据源 → 添加数据源 → 钉钉文档**,填写凭证并测试连接,
然后选择同步范围和同步策略。

> 当前仅同步节点类型为 `FILE`、类别为 `ALIDOC`、扩展名为 `adoc` 的钉钉在线文档。
> 在线表格、多维表及普通上传文件不会被错误地当作文档块读取。只有目录遍历完整时才会
> 生成删除标记;分支读取失败会保留旧游标,避免临时权限或网络问题造成误删除。

---

## 前端管理
Expand Down
47 changes: 47 additions & 0 deletions frontend/src/i18n/datasourceConnectorLocale.test.ts
Original file line number Diff line number Diff line change
@@ -0,0 +1,47 @@
import assert from 'node:assert/strict'
import { readFileSync } from 'node:fs'
import { dirname, join } from 'node:path'
import { test } from 'node:test'
import { fileURLToPath } from 'node:url'

import { LOCALE_BUNDLES, getLocaleValueAtPath, type LocaleName } from './localeKeyAudit.ts'

const DIALOG_PATH = join(
dirname(fileURLToPath(import.meta.url)),
'../views/knowledge/settings/DataSourceEditorDialog.vue',
)

// The connector picker resolves its label and description through computed keys
// (`datasource.connector.${def.type}`), which the static usage audit cannot
// see. A type declared in connectorDefs but absent from a locale therefore
// renders the raw key in the UI while every other check stays green, so assert
// the two bags directly against the connector list that drives the picker.
function declaredConnectorTypes(): string[] {
const source = readFileSync(DIALOG_PATH, 'utf8')
const start = source.indexOf('const connectorDefs')
assert.notEqual(start, -1, 'connectorDefs not found in DataSourceEditorDialog.vue')
const end = source.indexOf('const currentDef', start)
assert.notEqual(end, -1, 'end of connectorDefs not found')
const types = [...source.slice(start, end).matchAll(/^\s*type:\s*'([^']+)'/gm)].map((m) => m[1])
assert.ok(types.length > 0, 'no connector types parsed from connectorDefs')
return types
}

test('every connector type has a name and description in every locale', () => {
const types = declaredConnectorTypes()
const failures: string[] = []

for (const type of types) {
for (const bag of ['connector', 'connectorDesc'] as const) {
for (const [localeName, bundle] of Object.entries(LOCALE_BUNDLES) as Array<
[LocaleName, unknown]
>) {
const path = `datasource.${bag}.${type}`
const label = getLocaleValueAtPath(bundle, path)
if (typeof label !== 'string' || !label) failures.push(`${localeName}: missing ${path}`)
}
}
}

assert.deepEqual(failures, [], failures.join('\n'))
})
12 changes: 12 additions & 0 deletions frontend/src/i18n/embed.ts
Original file line number Diff line number Diff line change
Expand Up @@ -106,6 +106,9 @@ const messages = {
"unableToGetKnowledgeBaseId": "无法获取知识库ID",
"summaryInProgress": "正在总结答案……",
"thinkingAlt": "正在思考",
"preparingAnswer": "正在准备回答…",
"connectingModelAndGeneratingAnswer": "正在连接模型并生成回答…",
"modelStillResponding": "模型响应较慢,仍在等待…",
"deepThoughtCompleted": "已深度思考",
"deepThoughtAlt": "深度思考完成",
"referencesTitle": "参考了{count}个相关内容",
Expand Down Expand Up @@ -596,6 +599,9 @@ const messages = {
"unableToGetKnowledgeBaseId": "Unable to get knowledge base ID",
"summaryInProgress": "Summarizing answer…",
"thinkingAlt": "Thinking in progress",
"preparingAnswer": "Preparing an answer…",
"connectingModelAndGeneratingAnswer": "Connecting to the model and generating an answer…",
"modelStillResponding": "The model is taking longer than usual, still waiting…",
"deepThoughtCompleted": "Deep thinking completed",
"deepThoughtAlt": "Deep thinking finished",
"referencesTitle": "Referenced {count} related item(s)",
Expand Down Expand Up @@ -1048,6 +1054,9 @@ const koEmbedPublish = {
followUpQuestions: '이어서 질문',
followUpQuestionsLoading: '추천 질문 로딩 중',
thinkingAlt: '생각 중',
preparingAnswer: '답변을 준비하고 있습니다…',
connectingModelAndGeneratingAnswer: '모델에 연결하여 답변을 생성하고 있습니다…',
modelStillResponding: '모델 응답이 평소보다 오래 걸리고 있습니다. 계속 기다리는 중…',
refreshSuggestedQuestions: '다른 질문',
imageTooMany: '이미지는 최대 5장까지 업로드할 수 있습니다',
imageTypeSizeError: 'JPG/PNG/GIF/WEBP만 지원하며, 각 파일은 10MB 이하여야 합니다',
Expand Down Expand Up @@ -1136,6 +1145,9 @@ const ruEmbedPublish = {
followUpQuestions: 'Спрашивайте дальше',
followUpQuestionsLoading: 'Загрузка рекомендуемых вопросов',
thinkingAlt: 'Обдумывание...',
preparingAnswer: 'Подготовка ответа…',
connectingModelAndGeneratingAnswer: 'Подключение к модели и создание ответа…',
modelStillResponding: 'Модель отвечает дольше обычного, продолжаем ждать…',
refreshSuggestedQuestions: 'Ещё',
imageTooMany: 'Можно загрузить не более 5 изображений',
imageTypeSizeError: 'Поддерживаются только JPG/PNG/GIF/WEBP, каждый файл до 10 МБ',
Expand Down
Loading