Skip to content
151 changes: 151 additions & 0 deletions docs/插件开发指南.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,151 @@
# 插件开发指南

WeKnora 的可扩展子系统统一建立在插件内核 `internal/plugin` 之上。本文说明如何为已有领域贡献一个插件、如何把一个新子系统改造成可插拔的领域,以及进程外插件的接入方式。

## 内核解决什么问题

在内核出现之前,仓库里已经长出了四套互不相通的注册表——联网搜索、文档解析引擎、数据源连接器、模型厂商。它们各自回答同一组问题,而且答案不一致:

| 问题 | 内核之前 | 内核之后 |
|---|---|---|
| 这是什么,能干什么 | 有的只有 id,有的有 Name/Description | `Manifest` |
| 需要什么配置 | 有的用 Go struct,有的用 `map[string]string` 魔法键 | `Schema` |
| 配置是否合法 | 各自在构造函数里手写 if | 内核在构造前校验 |
| 现在能不能用,不能用为什么 | 只有文档解析回答了,其余靠猜 | `Health` |
| 前端表单怎么渲染 | 单独维护一份列表,容易漂移 | `Schema.Form()` |

四份答案意味着四倍的维护面,而且实际上已经漂移了:模型层的规则被手工复制到了前端。内核把这五个答案统一,代价是每个插件多写一段声明。

## 核心概念

```go
type Plugin[T any] interface {
Manifest() Manifest // 身份与能力
Schema() Schema // 配置契约
Probe(ctx, Config) Health // 现在能不能用
New(ctx, Config) (T, error) // 造一个实例
}
```

`T` 是**领域自己定义的能力接口**,内核不认识它。这样 `Registry[WebSearchProvider]` 不可能返回一个聊天模型,而跨领域的统一列表由非泛型的 `Catalog` 提供。

三条设计约束,写插件时请遵守:

1. **只声明事实,不分支于名字。** 任何需要 `switch pluginID` 的东西,都应该变成描述符里的一个字段。
2. **配置的唯一契约是 Schema。** 不要读 Schema 没声明的键;需要新选项就加字段,它会自动获得校验和表单控件。
3. **可用性是上报的事实,不是承诺。** `Probe` 返回 `Degraded` 时,需要完整能力的调用方必须视为拒绝,而不是四舍五入成可用。

## 给已有领域加一个插件

以联网搜索为例。完整例子见 `internal/infrastructure/web_search/plugins.go`。

### 1. 实现能力接口

领域已经定义好了接口(这里是 `interfaces.WebSearchProvider`),照常实现即可。构造函数保持原样——**不需要**在里面校验配置,那是 Schema 的事。

```go
func NewAcmeProvider(params types.WebSearchProviderParameters) (interfaces.WebSearchProvider, error) {
client, err := NewSearchHTTPClient(15*time.Second, params.ProxyURL)
if err != nil {
return nil, err
}
return &AcmeProvider{client: client, apiKey: params.APIKey}, nil
}
```

### 2. 声明它

```go
func init() {
Register(Definition{
ID: "acme",
DisplayName: "Acme Search",
DocURL: "https://acme.example/docs/api",
Fields: []Field{
APIKeyField(true),
OptionField(FieldScope, 10, "web", "web", "news"),
},
New: NewAcmeProvider,
})
}
```

到此为止你已经获得:配置校验、设置页表单、`/plugins` 目录里的一条记录、以及统一的错误形态。**不需要**去改依赖注入容器,也不需要去改前端。

### 3. 需要 Schema 表达不了的检查时

Schema 能表达"必填""在这个词表里""在这个数值区间"。表达不了的(比如"这个 URL 必须是绝对的 http(s) 且不指向内网")放进 `Validate`,它会在 `Probe` 时运行:

```go
Validate: func(params Params) error { return ValidateAcmeEndpoint(params.BaseURL) },
```

注意 `Probe` 可能做真实 I/O(SearXNG 的校验会做 DNS 解析)。这正是 `Probe` 与 Schema 校验分开的原因:保存表单前想快速校验就只跑 Schema,想知道"真的能连上吗"才跑 Probe。

### 4. 写测试

测行为,不测实现。参考 `internal/infrastructure/web_search/plugin_test.go`:缺凭据要在构造函数之前被拒、选项的错误取值要在本地被拒而不是发给上游、密钥不能通过目录漏出去。

## 把一个新子系统改造成可插拔领域

需要三样东西,加起来通常不到 100 行。

### 1. 定义能力接口和 Kind

```go
// Kind 用点号分域,UI 可以按前缀分组而不必认识每一个。
const Kind = plugin.Kind("chunking")

type Chunker interface {
Chunk(ctx context.Context, doc []byte) ([]Chunk, error)
}

var Plugins = plugin.NewRegistry[Chunker](Kind)
```

### 2. 给领域一个便捷的声明入口

不是必须的,但能让插件声明读起来干净。web_search 的 `Definition` + `Register` 就是这一层:它把「已有的构造函数」适配成内核的 `Plugin[T]`,顺便自动附上每个插件都需要的公共字段(比如代理地址)。

### 3. 让消费方走 `Open`

```go
chunker, err := chunking.Plugins.Open(ctx, cfg.Strategy, cfg.Params)
```

`Open` 会先按 Schema 校验再构造。绕过它自己 `Lookup` + 手工构造,等于放弃了这个保证——所以 `Open` 是文档化的入口。

### 迁移既有子系统的注意事项

- **存储格式不用动。** 写一对 `RawFromParams` / `ParamsFromConfig` 的桥接函数,旧配置无需数据迁移就能进入新路径。web_search 就是这么做的。
- **实现不用动。** 只改注册方式,构造函数原样保留,回归面最小。
- **迁移完要删掉旧注册表**,否则两套并存比一套混乱更糟。

## 进程外插件

Go 没有好用的动态加载,所以第三方插件走两条路:

**嵌入式**:把 WeKnora 当库用的场景,直接调 `Registry.Register` 注册自己的实现,无需 fork。注册是可逆的,返回的函数就是注销。

**远程插件**:`Manifest` 和 `Schema` 都是可序列化的,所以一个跑在别处的插件可以发布同样的清单,由 `Catalog.PublishExternal` 登记进目录:

```go
undo, err := plugin.DefaultCatalog.PublishExternal(manifest, schema)
```

外部插件在目录里带 `external: true` 标记,有清单和表单但没有进程内工厂,由所属领域决定怎么驱动它(HTTP、gRPC、子进程都可以)。文档解析领域已经在通过 RPC 发现远程引擎并和本地引擎合并,那套手写的合并逻辑正是这个机制要取代的。

## 前端如何消费

不要在前端复刻后端规则。`Catalog` 已经把每个插件的表单结构(分组、字段、控件类型、词表、i18n key)暴露出来,前端按结构渲染即可。

后端只发**结构和 i18n key**,不发显示文本——语言归前端管。这条约束的由来是一个真实教训:模型层的厂商规则曾被复制到 `frontend/src/utils/thinkingControl.ts`,注释里写着"必须与后端保持一致",然后它们就不一致了。一句请求后人保持同步的注释不是机制。

## 参考实现

| 位置 | 看什么 |
|---|---|
| `internal/plugin/` | 内核本身;测试用虚构领域,证明它不认识任何具体业务 |
| `internal/infrastructure/web_search/plugin.go` | 领域适配层:Definition、公共字段、存储桥接 |
| `internal/infrastructure/web_search/plugins.go` | 插件声明集合 |
| `internal/infrastructure/web_search/plugin_test.go` | 该测什么 |
116 changes: 116 additions & 0 deletions docs/模型插件化设计.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,116 @@
# 模型管理插件化设计

本文是模型域迁移到插件内核(`internal/plugin`)的设计。它记录**为什么第一版不合格**、目标形态、以及要删掉什么。

## 第一版为什么不合格

第一版做了一个干净的模型能力缝,然后把它螺栓到了旧结构上。结果是旧的混乱一个都没消失,只是多了一层新抽象盖在上面:

| 旧的混乱 | 第一版的处理 | 问题 |
|---|---|---|
| `provider.DetectProvider(baseURL)` 按 URL 嗅探厂商 | 继续调用它 | 用户无法可靠指定厂商;一个 URL 改写就换了一套参数语义 |
| `ChatConfig` 上帝结构,混进了 WeKnoraCloud 专用的 AppID/AppSecret | 继续用它 | 每加一个需要特殊凭据的厂商,这个结构就长一个字段 |
| `extra_config map[string]string` 魔法键 | 继续用它,还新增了 `protocol` 键 | 键名只在代码里,没有契约 |
| `RemoteAPIChat` 的 SDK/raw 双路径 | 保留,用报文 diff 决定走哪条 | 一个请求两种可能的传输,行为难以推理 |
| `providerAdapter` 传输层特判 | 保留 | 和插件描述符两套厂商概念并存 |
| `ModelSource` 枚举与真实路由无关 | 未处理 | 误导性字段仍在 |

而且那层新抽象**只服务模型**:配置、值系统、表单渲染这些明显通用的东西被放在了 `internal/models/llm/spi` 里,别的子系统用不上。这是位置放错,不是抽象错。

## 目标形态

三层,关键是**内核不知道什么是模型**。

```
internal/plugin/ 领域中立的内核(已完成)
Manifest / Schema / Health / Registry[T] / Catalog

internal/models/llm/ LLM 领域,建在内核上
wire/ Encoder、Draft、Plan —— 只有 HTTP 模型 API 才有的"参数→报文字段"映射
protocol/ 三个标准协议驱动(已完成,无需改动)
vendors/ 厂商插件声明
runtime/ 自己的客户端:解析插件 → 构造报文 → 发请求 → 解码

internal/models/chat/ 只剩一层薄适配器,把旧接口桥到新运行时
```

### 领域能力接口

```go
const KindChat = plugin.Kind("llm.chat")

// Chat 是领域自己的能力接口,内核不认识它。
var ChatPlugins = plugin.NewRegistry[Chat](KindChat)
```

`llm.embedding`、`llm.rerank`、`llm.vision`、`llm.asr` 是同构的兄弟 Kind,各自一个注册表。这直接解决了 embedding / rerank 现在那两个 giant switch。

### 配置:用内核的 Schema,不再自建

第一版在 `llm/spi` 里自己写了一套 `Value` / `Param` / `FieldSchema`。这些**全部删除**,改用 `plugin.Value` / `plugin.Field` / `plugin.Schema`。LLM 层只保留内核表达不了的东西:

```go
// Param 是一个内核字段,外加"它写到报文的哪个位置"。
// 这是 LLM 域独有的概念:内核管配置,wire 管报文。
type Param struct {
plugin.Field
Encode Encoder // enable_thinking / thinking.type / chat_template_kwargs / ...
}
```

这样 `thinking.mode` 既是一个有值域、有表单控件的配置字段(内核负责),又知道自己落到哪个 JSON 字段(LLM 域负责)。

### 模型规格:显式,不嗅探

`ChatConfig` 被替换为:

```go
type ModelSpec struct {
ModelID string // WeKnora 内部 id
PluginID string // 显式的厂商插件 id,绝不从 URL 推断
Model string // 发给厂商的模型名
Protocol plugin.Kind // 可选,厂商支持多协议时才需要
Config map[string]any // 交给插件 Schema 校验
}
```

- **没有 `Source` 枚举**:本地 Ollama 就是一个 plugin id 为 `ollama` 的插件。
- **没有 AppID/AppSecret 专用字段**:需要应用级凭据的厂商在自己的 Schema 里声明两个 `Secret` 字段。
- **没有 `extra_config`**:所有配置都是 Schema 声明过的字段。
- **没有 URL 嗅探**:`PluginID` 为空时不猜,报错要求用户选择。

### 迁移期的兼容

存量数据里 `PluginID` 是空的、配置在 `parameters` 和 `extra_config` 里。处理方式和 web_search 一样:写一对桥接函数,**不做数据迁移**。

```go
// 只在读取存量记录时调用一次;写入路径永远写显式 PluginID。
func SpecFromLegacyModel(m *types.Model) ModelSpec
```

其中对空 `PluginID` 的记录,**保留一次性的 URL 嗅探作为兜底**,并在日志和 API 响应里标记为"推断得到",提示用户去设置页确认。嗅探从此只存在于这一个函数里,而不是散布在请求路径上。

## 要删掉什么

迁移完成后这些应当消失,而不是共存:

- `internal/models/chat/provider.go` 的 `providerAdapter` 体系 —— 它剩下的四项传输行为(签名、非派生 endpoint、消息改写、工具调用元数据)改为插件描述符字段或插件自己的 `New` 实现。
- `internal/models/chat/remote_api.go` 的 SDK/raw 双路径 —— 新运行时始终自己序列化,只有一条传输路径。
- `internal/models/llm/spi` 里的 `Value` / `Param` / `ParamUI` / `FieldSchema` / `Registry` —— 由内核提供。
- `ChatConfig`、`ExtraConfigThinkingControl`、`ExtraConfigProtocol`。
- `provider.DetectProvider` 在请求路径上的所有调用点(只保留在存量桥接里)。
- embedding / rerank / vlm / asr 的 `switch` 工厂。

## 迁移顺序

每一步都可独立合入、可独立回滚:

1. **wire 层归位**:`llm/spi` 的配置概念换成内核的,只留 `Encoder` / `Draft` / `Plan`。厂商声明和黄金报文测试基本不动。
2. **chat 领域上内核**:定义 `KindChat` 与 `ChatPlugins`,厂商描述符改为 `plugin.Plugin[Chat]`,新运行时实现三协议的完整路径。
3. **切换调用方**:`chat.NewChat` 变成薄适配器,内部走新运行时;旧的 `RemoteAPIChat` 与 `providerAdapter` 删除。
4. **其余四种模型类型**:embedding / rerank / vision / asr 各自一个 Kind,消灭对应的 switch。
5. **前端**:模型编辑表单改为按 `Catalog` 的表单结构通用渲染,不再为模型单独写一套。

## 为什么不一次做完

第 3 步会改动模型调用链上的每一个消费方(RAG 问答、Agent、抽取、摘要、向量化)。分步做是为了让每一步的回归面都可控,而不是一次性替换后再去查是哪里坏了。
68 changes: 68 additions & 0 deletions frontend/src/api/model/index.ts
Original file line number Diff line number Diff line change
Expand Up @@ -259,3 +259,71 @@ export function getWeKnoraCloudStatus(): Promise<WeKnoraCloudStatusResult> {
})
})
}

// ---------------------------------------------------------------------------
// Model capability manifest
// ---------------------------------------------------------------------------

import type { ModelCapabilities } from '@/utils/modelCapabilities'

export interface CapabilityQuery {
provider?: string;
model?: string;
baseUrl?: string;
modelType?: string;
protocol?: string;
}

function capabilityCacheKey(query: CapabilityQuery): string {
return [
query.provider ?? '',
query.model ?? '',
query.baseUrl ?? '',
query.modelType ?? 'chat',
query.protocol ?? '',
].join('|');
}

// Capabilities are a pure function of the query on the server, so caching them
// avoids a request per keystroke while a user types a model name.
const capabilityCache = new Map<string, Promise<ModelCapabilities | null>>();

/**
* Fetch the capability manifest for a provider and model.
*
* Resolves to null when the backend has no plugin for the provider, which is a
* gap in the catalog rather than an error: the model still works through the
* generic transport, and the caller should fall back to its own defaults.
*/
export function fetchModelCapabilities(query: CapabilityQuery): Promise<ModelCapabilities | null> {
if (!query.provider && !query.baseUrl) {
return Promise.resolve(null);
}

const key = capabilityCacheKey(query);
const cached = capabilityCache.get(key);
if (cached) return cached;

const params: Record<string, string> = { model_type: query.modelType ?? 'chat' };
if (query.provider) params.provider = query.provider;
if (query.model) params.model = query.model;
if (query.baseUrl) params.base_url = query.baseUrl;
if (query.protocol) params.protocol = query.protocol;

const request = get('/models/capabilities', params)
.then((res: any) => (res?.data ?? null) as ModelCapabilities | null)
.catch(() => {
// A failed lookup must not block the editor; drop it so a later attempt
// can succeed.
capabilityCache.delete(key);
return null;
});

capabilityCache.set(key, request);
return request;
}

/** Clear the capability cache, for an explicit refresh. */
export function clearCapabilityCache(): void {
capabilityCache.clear();
}
36 changes: 33 additions & 3 deletions frontend/src/components/ModelDebugDrawer.vue
Original file line number Diff line number Diff line change
Expand Up @@ -197,9 +197,9 @@ import { MessagePlugin } from 'tdesign-vue-next'
import { useI18n } from 'vue-i18n'
import { copyWithToast } from '@/utils/clipboard'
import SettingDrawer from '@/components/settings/SettingDrawer.vue'
import { debugModel, type ModelConfig, type ModelDebugResult } from '@/api/model'
import { debugModel, fetchModelCapabilities, type ModelConfig, type ModelDebugResult } from '@/api/model'
import { fileSizeVerification } from '@/utils'
import { modelSupportsThinking } from '@/utils/thinkingControl'
import { supportsThinking as capabilitiesSupportThinking } from '@/utils/modelCapabilities'

const props = defineProps<{
visible: boolean
Expand Down Expand Up @@ -242,7 +242,37 @@ let runSequence = 0
const selectedModel = computed(() => props.models.find(model => model.id === selectedModelId.value))
const filteredModels = computed(() => props.models.filter(model => model.type === selectedModelType.value))
const isChat = computed(() => selectedModel.value?.type === 'KnowledgeQA')
const supportsThinking = computed(() => selectedModel.value ? modelSupportsThinking(selectedModel.value) : false)
/**
* Ask the backend whether a saved model has a thinking toggle. Only remote
* chat models can have one, and a stored legacy override still wins because
* the backend still honors it.
*/
async function resolveSupportsThinking(model: ModelConfig): Promise<boolean> {
if (model.type !== 'KnowledgeQA' || model.source !== 'remote') return false
const capabilities = await fetchModelCapabilities({
provider: model.parameters.provider || '',
model: model.name || '',
baseUrl: model.parameters.base_url || '',
})
return capabilitiesSupportThinking(
capabilities,
model.parameters.extra_config?.thinking_control,
)
}

// Whether the selected model has a thinking toggle is the backend's answer,
// resolved from the model plugin, so the drawer cannot offer a switch the
// model will ignore. It arrives asynchronously, hence a ref rather than a
// computed.
const supportsThinking = ref(false)
watch(
selectedModel,
async (model) => {
supportsThinking.value = model ? await resolveSupportsThinking(model) : false
if (!supportsThinking.value) thinking.value = false
},
{ immediate: true },
)
const needsFile = computed(() => ['VLLM', 'ASR'].includes(selectedModel.value?.type || ''))
const documents = computed(() => documentsText.value.split('\n').map(item => item.trim()).filter(Boolean))
const canRun = computed(() => {
Expand Down
Loading