Files
2026-02-02 11:22:35 +08:00

185 lines
9.1 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# User Profile Scoring技术计划 · V1.2
## 1. 计划目标
在**客户端TypeScript**实现一个可复用的“用户画像计算”模块:
- **输入**Onboarding 问卷答案(每题**可跳过**;字段可缺失/为 `null`
- **输出**:严格符合 `spec_kit/User Profile Scoring/spec.md`V1.2)的用户画像对象(含 `profile_answered`、四维度字段、`profile_confidence`
- **硬规则**:输出一份“下游可直接执行/可观测”的硬规则结果,用于推荐/Push/Widget 统一复用(对齐 V1.2 规则语义)
- **约束**:不抛异常、不报错;非法值按**跳过**处理
## 2. 默认技术决策(本计划采用)
- **实现位置**`client/src/features/userProfileScoring/`
- **语言**TypeScript纯函数为主便于测试与跨端复用
- **时间与可测性**:允许调用方注入 `generatedAt/now`,默认使用 `new Date()`,保证可回归测试
- **规则对齐来源**
- 画像与置信度:`设计说明文档/客戶端問卷打分規則.md`V1.2
- 下游过滤语义:`设计说明文档/個性化推薦算法規則.md`Hard Filter / `risk_flags` 语义)
## 3. 输入与输出契约V1.2,含跳过)
### 3.1 输入:问卷答案(允许部分缺失/跳过)
定义输入类型 `QuestionnaireAnswersV1_2`(所有字段可选;也允许显式 `null`
- `mom_stage?: "expecting" | "parenting" | "unknown" | null`
- `emotion?: "low" | "overwhelmed" | "tired" | "neutral" | "calm" | "joyful" | null`
- `context?: "family" | "work" | "relationship" | "friends" | "health" | null`
- `need?: "emotional_support" | "parenting_pressure" | "self_worth" | "anxiety_relief" | "rest_balance" | null`
非法值处理策略:
- 字段存在但值不在枚举内:按**跳过**处理(归一化为 `undefined`),不抛错
### 3.2 输出:用户画像(严格对齐 spec.md V1.2
输出 `UserProfileV1_2`,满足 `spec.md` 的字段与取值:
- `profile_version: "v1.2"`
- `profile_source: "questionnaire"`
- `profile_generated_at: string`ISO8601
- `profile_confidence: number`01
- `profile_answered: { stage: boolean; emotion: boolean; context: boolean; need: boolean }`
- `stage: { expecting?: 0|1; parenting?: 0|1; unknown: 0|1 }`
- 注:题目跳过时按安全策略输出 `unknown=1`
- `emotion_score: number | null`
- `context: Record<string, 1>`(稀疏 one-hot题目跳过时为 `{}`
- `need: Record<string, 1>`(稀疏 one-hot题目跳过时为 `{}`
### 3.3 额外输出:硬规则(供下游直接执行/打点)
`spec.md` 要求“向下游暴露可直接执行的硬规则结果”,但不强制字段结构。本计划采用以下扩展字段(不破坏 `spec.md` 核心结构):
- `rule_hits: string[]`
- 仅用于可观测/回归测试(例如 `["unsafe_for_emotion_low"]`
- `hard_rules: { ... }`
- `forbidden_risk_flags: string[]`(对齐内容侧/推荐侧 `risk_flags` 名称)
- `forbidden_content_predicates: Array<{ id: string; when_user: object; forbid_content: object }>`
- 用于表达“需要同时看用户与内容字段才能执行”的规则(例如 unknown + 育儿压力 + 强个性化)
## 4. 跳过题目的处理策略V1.2:安全优先 + 可观测)
> 目标:题目跳过时不报错;画像仍可计算;并通过 `conf_U` 与下游 `P_uncertainty` 思路自动降个性化/降风险。
- **mom_stage 跳过**
- 输出:`stage.unknown = 1`
- `profile_answered.stage = false`
- **emotion 跳过**
- 输出:`emotion_score = null`(不强行填 0.6,避免制造“伪精确画像”)
- `profile_answered.emotion = false`
- **context 跳过**
- 输出:`context = {}`
- `profile_answered.context = false`
- **need 跳过**
- 输出:`need = {}`
- `profile_answered.need = false`
## 5. 画像计算规则(映射 + 置信度 + 硬规则)
### 5.1 映射规则(对齐 `客戶端問卷打分規則.md` V1.2
- `mom_stage``stage.*`one-hot若跳过则 `unknown=1`
- `emotion``emotion_score`
- `low=0.0``overwhelmed=0.2``tired=0.4``neutral=0.6``calm=0.8``joyful=1.0`
- 若跳过:`emotion_score = null`
- `context` → 稀疏 one-hot例如 `work``{ work: 1 }`;若跳过:`{}`
- `need` → 稀疏 one-hot例如 `rest_balance``{ rest_balance: 1 }`;若跳过:`{}`
### 5.2 置信度(`profile_confidence / conf_U`计算V1.2
时间衰减(与 `spec.md` 一致):
- 07 天:`conf_time = 1.0`
- 730 天:线性衰减到 `0.7`
- 30 天以上:`conf_time = 0.5`
作答完整度因子(对齐 `客戶端問卷打分規則.md` V1.2
- `answered_count = Σ I[profile_answered.* = true]``total_questions = 4`
- `completion = answered_count / total_questions`
- `completion_factor = 0.5 + 0.5 * completion`
- `profile_confidence = clamp(conf_time * completion_factor, 0.2, 1.0)`
### 5.3 硬规则输出(对齐 V1.2
本模块输出 `rule_hits``hard_rules`,用于下游先过滤、后打分:
1) **按用户状态映射到内容 `risk_flags` 的禁推**
-`stage.unknown = 1`
- `rule_hits += ["unsafe_for_stage_unknown"]`
- `hard_rules.forbidden_risk_flags += ["unsafe_for_stage_unknown"]`
-`stage.parenting = 1`
- `rule_hits += ["unsafe_for_stage_parenting"]`
- `hard_rules.forbidden_risk_flags += ["unsafe_for_stage_parenting"]`
-`emotion_score !== null``emotion_score <= 0.2`
- `rule_hits += ["unsafe_for_emotion_low"]`
- `hard_rules.forbidden_risk_flags += ["unsafe_for_emotion_low"]`
2) **跨维度产品规则(需要同时看用户与内容字段)**
-`stage.unknown = 1`
- 增加一个可执行谓词,表达“禁推 育儿压力且高个性化内容”:
- `when_user`: `{ stage_unknown: true }`
- `forbid_content`: `{ need: "parenting_pressure", personalization_power: 1 }`
- 该规则由下游在内容侧字段存在时执行(对齐 `spec.md` 7 与 `客戶端問卷打分規則.md` 6 的 V1.2 口径)
> 注:本模块不判断内容的 `personalization_power`(属于内容侧),只输出“可执行条件”。
## 6. 文件与目录规划(客户端)
`client/src/` 新增/调整:
- `client/src/features/userProfileScoring/types.ts`
- `QuestionnaireAnswersV1_2` / `UserProfileV1_2` / 枚举联合类型
- `client/src/features/userProfileScoring/scoring.ts`
- `buildUserProfileFromQuestionnaire(answers, options): UserProfileV1_2 & { rule_hits; hard_rules }`
- `computeTimeConfidence(generatedAt, now): number`
- `normalizeAnswers(raw): QuestionnaireAnswersV1_2`(非法值 → 跳过)
- `computeProfileAnswered(normalized): profile_answered`
- `computeProfileConfidence(conf_time, profile_answered): number`
- `client/src/features/userProfileScoring/index.ts`
- 对外导出(供 Onboarding/推荐/Push 调用)
## 7. 测试计划(客户端可回归)
若工程未提供 `test` 脚本,建议补齐最小单元测试能力(只测纯函数):
- 新增 `vitest`(或 `jest`)为 devDependencies并在 `client/package.json` 增加 `test` 脚本
-`buildUserProfileFromQuestionnaire` 增加用例(与 `spec.md` 第 10 章一致):
- 映射正确性(完整作答)
- 置信度时间衰减07/730/30+
- 硬规则触发(`emotion_score<=0.2``stage.unknown``stage.parenting`
- 跳过场景V1.2
- 全部跳过(输入 `{}``profile_answered` 全 false`stage.unknown=1``emotion_score=null``context/need={}``profile_confidence=0.5`(当 conf_time=1 时)
- 仅回答 1 题:验证 `completion_factor` 与 clamp最小 0.2)逻辑正确
## 8. 实施步骤
1.`client/src/features/userProfileScoring/` 增加类型定义与常量映射表V1.2
2. 实现答案归一化(非法值按跳过;支持 `null/undefined`
3. 实现画像生成主函数(含 `profile_answered`、元信息生成、置信度计算)
4. 实现硬规则命中与结构化输出(`forbidden_risk_flags` + `forbidden_content_predicates`
5. 补齐最小测试框架与单测(仅纯函数)
6. Onboarding 流程完成处接入(后续任务):问卷提交后生成画像对象;持久化/上传由上层决定
## 9. 验收标准
- 任意问卷答案输入(含任意题跳过/非法值),都能生成符合 `spec.md`V1.2)结构的用户画像对象(不抛错)
- 输出包含:
- V1.2 元信息(`profile_version/source/generated_at/confidence/profile_answered`
- 四维度字段(`stage/emotion_score/context/need`;允许跳过后的 `null/{}`
- `rule_hits``hard_rules`(供下游执行 Hard Filter 与打点)
- 置信度:
- 时间衰减符合 `spec.md` 6.2
- 跳过越多 `profile_confidence` 越低(按 V1.2 完整度因子),并被 clamp 到 `[0.2, 1.0]`
## 10. 后续演进(不阻塞 V1.2
- 支持 `context/need` 多选(输出仍为 one-hot可多个 key=1
- 引入 `profile_source="mixed"`:行为信号修正 `emotion_score` 或补全缺失维度(保留 questionnaire 原始画像用于可观测)
- 将硬规则输出进一步与推荐系统规则引擎 schema 对齐(使 `forbidden_content_predicates` 可直接落入规则引擎)