Palantir 端到端实战:工厂 IoT 告警自动分派
以「工厂 IoT 告警自动分派」为场景,把 Ontology → Action → Automate → AIP Agent 串成一条端到端链路。同一套骨架套用到供应链缺料、风控拦截、运维排障等场景同样成立,只换 Object Type 和参数即可。
场景设定
- 业务诉求:传感器上报温度异常 → 系统自动建 Alert 对象 → 按温度区间和工厂位置决定优先级与处理组 → 自动填字段、建 Link、发通知 → 严重告警让 AIP Agent 辅助生成处理建议,但不直接改物理设备状态(需人确认)。
一、Ontology 层:先定义”名词”
以下 Object Type 均建模在 Foundry Ontology 中:
Object Type:Alert(告警)
| 属性 | 类型 | 说明 |
|---|---|---|
| id | string | 告警唯一标识 |
| temperature | float | 当前温度读数 |
| detectedAt | datetime | 告警检测时间 |
| priority | enum | LOW / MEDIUM / HIGH / CRITICAL |
| status | enum | OPEN / ACK / CLOSED |
| factoryId | string | 所属工厂 |
Links:
assignedTo→ Group(处理组,分派后建立)originatesFrom→ Factory(多对一)
Object Type:Factory(工厂)
| 属性 | 类型 |
|---|---|
| factoryId | string |
| region | string |
| name | string |
Object Type:Group(处理组)
| 属性 | 类型 | |
|---|---|---|
| groupId | string | |
| name | string | |
| scope | enum | HVAC / 电气 / 安全 |
关键点
Action 后面操作的是这些语义对象,而非直接写底层 IoT 数据表。LLM/Agent 看到的是”告警、工厂、处理组”,不接触原始存储。
二、Action 层:定义”动词”(受治理的写操作)
2.1 Action Type:AssignAlertToGroup
将告警分派给对应处理组,是一次事务性写操作。
Parameters
| 参数 | 类型 | 必填 | 说明 |
|---|---|---|---|
| alert | Alert | ✅ | 待分派的告警对象 |
| group | Group | ✅ | 接收分派的组 |
| priority | enum | ❌ | 默认由规则计算,也可显式传入 |
Edits(提交时对 Ontology 做事务写)
- 修改
Alert.priority - 建立 Link
Alert.assignedTo → Group - 若原
status = OPEN,改为ACK
Submission Criteria(事前门禁)
- Alert.status 必须为 OPEN
- 调用者须具备
alert-management角色 - Group.scope 须与 Alert 来源 Factory.region 匹配(复杂规则用 Function-backed 判断)
Permissions
- RBAC:仅 SRE / 厂务值班组可直接提交
- Marking:含 PII 标记的 Factory 数据,Agent 无对应 Marking 时看不到也不能写
Side Effects(成功后触发)
- 发 Foundry 站内通知给 Group 成员:”Alert X 已分派给你”
- Webhook → 企业微信/钉钉机器人播报
Idempotency & Audit
- 同一 Alert 重复提交:若已 ACK 且 Link 存在,平台视为 no-op 或返回既有状态
- 每次执行写 Action Execution 表:谁、何时、什么参数、成功/失败
2.2 Action Type:UpdateAlertStatus
供 Agent 或人工调用,用于更新告警状态。
- 参数:
alert、newStatus - 编辑:改
Alert.status - 权限:Agent 可调用;但
CLOSED状态要求人工二次确认(Human-in-the-loop) - 审计:记录 LLM 决策来源(Provenance 标为 AIP Agent)
三、Automate 层:事件驱动编排(条件 → 效果)
在 Automate 应用中创建一个 Rule:
Trigger(条件)
- 类型:Object 数据条件
- 监听:Alert 对象创建
- 过滤:
Alert.status = OPEN且Alert.temperature > 80
Effects(按顺序)
- Function 计算(TypeScript / Python)
- 输入:Alert.temperature、Factory.region
- 输出:
priority = temperature > 90 ? CRITICAL : HIGHgroup = 按 region + scope 映射出的 Group 对象
- Submit Action:AssignAlertToGroup
alert= 触发对象group= 上一步输出priority= 上一步输出
- 通知(可选,并行):给值班经理发邮件摘要
高级配置
| 能力 | 配置 |
|---|---|
| Retry | Action 提交失败自动重试 3 次(指数退避) |
| Fallback | Group 查不到时路由到 UnassignedAlertQueue 并告警值班 |
| 执行模式 | per-object(每新增一个 Alert 执行一次) |
部署后无需任何 cron 或外置调度——Foundry 监听 Ontology 变更自动触发。
四、AIP Agent 接入:让 LLM 在边界内辅助
值班人员问 AIP Agent:
“帮我看看 Alert #123,温度 95 度,给我处理建议,并按规则分派”
Agent 内部流程
1. Context Binding
将 #123 解析为具体 Alert 对象,自动加载 temperature、Factory、历史记录。
2. Reasoning(混合模式)
- 确定性代码:查阈值表 → 判定 CRITICAL,应分派给 HVAC 组
- LLM:基于 Alert 描述 + 知识库生成自然语言处理建议(如”检查冷却泵、通知厂务”)
3. Tool Call 校验
Agent 输出:AssignAlertToGroup(alert=123, group=HVAC-华东, priority=CRITICAL)
平台拦截检查:
- Agent 身份有无
alert-management权限? - Alert 是否处于 OPEN 状态?
- Region 是否匹配?
| 风险级别 | 处理方式 |
|---|---|
| 低风险(如更新为 ACK) | 自动通过 |
| 高风险(如直接 CLOSED、写设备指令) | 弹出人工确认卡(Human-in-the-loop) |
4. 执行 & 溯源
- 通过 → 提交 Action → Ontology 写回 + 钉钉通知
- 审计中
initiatedBy = AIP Agent,含 prompt 版本、function 版本
五、端到端时序
IoT 数据入 Foundry 原始表
→ Ontology 监听到新 Alert 对象
→ Automate 规则命中(temp > 80 且 OPEN)
→ 调 Function 算优先级 / 分组
→ 提交 AssignAlertToGroup Action(平台做权限 + 规则校验)
→ Ontology 更新(priority、Link、status)
→ Side Effect:钉钉通知 + Webhook
→ 值班人员问 AIP Agent
→ Agent 读 Ontology、跑建议、在治理边界内调 Action
→ 全部写 Action Execution + Lineage
六、这个设计为什么”企业级”
- AI 不直接碰库:Agent 只填 Action 参数,写操作统一走 Ontology Action
- 策略与代码分离:阈值、分派规则可放 Function 或单独 Rule 对象,业务改规则不用动 Agent prompt
- 完整审计:谁触发、LLM 哪次决策、哪条 Automate 规则、哪个 Action 提交,全链路可追溯
- 人机协同:低风险的自动跑,高风险的 Require Approval
提示:把本场景的 Object Type 替换为
Shipment、Inventory、RiskCase等,即可复用同一套骨架到供应链缺料自动调拨、风控拦截等场景。详见 Palantir 自动化机制概述。