个性化并不必然意味着“先认出这个人”。即使没有账号、用户画像或跨会话追踪 ID,系统仍然可以根据当前请求、用户主动选择的偏好以及本次会话中的行为,调整内容排序和界面呈现。
在没有更多来源实现细节的前提下,下面给出一种可落地的设计:客户端提交有限的即时上下文,服务端执行无状态打分,用完即弃,并明确禁止将这些信号用于权限和价格判断。
把“个性化”和“识别身份”拆开
传统个性化系统经常从一个稳定标识开始:账号 ID、广告 ID、设备 ID,或者长期 Cookie。随后,系统围绕这个标识积累浏览、购买和点击历史。
无身份个性化采用不同的数据流。它不问“这个人过去是谁”,而只处理完成当前任务所需的信号,例如:
- 用户在当前页面主动选择的语言、主题或内容类别;
- 当前时间段、页面入口和设备能力;
- 本次会话中刚刚浏览、隐藏或收藏的项目;
- 保存在设备本地、由用户控制的界面偏好;
- 不足以指向个人的粗粒度上下文,例如“移动端”而不是完整浏览器指纹。
这里的边界非常重要:匿名字段组合得足够多,仍可能形成指纹。IP 地址、精确位置、字体列表、屏幕尺寸、时区和行为时间戳叠加后,完全可能重新识别用户。因此,“没有 user_id 字段”不等于“没有身份风险”。
更稳妥的原则是只接受能够解释用途的字段:如果某个字段不能改变当前响应,就不要收集它。
一条无状态的数据路径
可以把请求设计成一次性的上下文计算:
客户端即时上下文
↓
字段校验与降精度
↓
无状态规则或模型打分
↓
返回排序结果
↓
丢弃原始上下文
客户端不需要发送账号或持久标识。例如,内容推荐接口可以只接收:
{
"locale": "zh-CN",
"period": "morning",
"recentCategories": ["developer"],
"dismissed": ["night-mode"]
}
这些信息描述的是“现在需要什么”,而不是“请求者是谁”。服务端也不必建立个人历史表。若要在页面切换间保留偏好,可以优先使用 sessionStorage;它的生命周期通常限于当前标签页会话。使用 localStorage 也不等同于建立账号,但会延长设备侧数据寿命,需要提供清除入口并说明用途。
可运行的 Node.js 无状态推荐接口
下面是一个只依赖 Node.js 标准库的最小示例。它不创建 Cookie、不接收用户 ID,也不保存请求正文。推荐结果只由本次提交的上下文计算。
将代码保存为 server.js,使用 Node.js 18 或更高版本运行:
const http = require("node:http");
const catalog = [
{
id: "docs-zh",
title: "中文开发文档",
categories: ["developer"],
periods: ["morning", "afternoon"],
locales: ["zh-CN"]
},
{
id: "api-playground",
title: "API 在线实验室",
categories: ["developer", "api"],
periods: ["afternoon", "evening"],
locales: ["*"]
},
{
id: "night-mode",
title: "夜间阅读模式",
categories: ["productivity"],
periods: ["evening", "night"],
locales: ["*"]
},
{
id: "getting-started",
title: "快速入门",
categories: ["beginner"],
periods: ["morning", "afternoon"],
locales: ["*"]
}
];
function normalizeContext(input) {
const allowedPeriods = new Set(["morning", "afternoon", "evening", "night"]);
return {
locale: typeof input.locale === "string"
? input.locale.slice(0, 16)
: "unknown",
period: allowedPeriods.has(input.period)
? input.period
: "unknown",
recentCategories: Array.isArray(input.recentCategories)
? input.recentCategories
.filter(value => typeof value === "string")
.slice(0, 5)
: [],
dismissed: Array.isArray(input.dismissed)
? input.dismissed
.filter(value => typeof value === "string")
.slice(0, 20)
: []
};
}
function recommend(context) {
const recent = new Set(context.recentCategories);
const dismissed = new Set(context.dismissed);
return catalog
.filter(item => !dismissed.has(item.id))
.map(item => {
let score = 0;
if (item.locales.includes(context.locale)) score += 3;
if (item.locales.includes("*")) score += 1;
if (item.categories.some(category => recent.has(category))) score += 2;
if (item.periods.includes(context.period)) score += 1;
return {
id: item.id,
title: item.title,
score
};
})
.sort((a, b) => b.score - a.score || a.id.localeCompare(b.id))
.slice(0, 3);
}
const server = http.createServer((req, res) => {
if (req.method !== "POST" || req.url !== "/recommend") {
res.writeHead(404, { "Content-Type": "application/json; charset=utf-8" });
return res.end(JSON.stringify({ error: "not_found" }));
}
let raw = "";
let tooLarge = false;
req.on("data", chunk => {
if (tooLarge) return;
raw += chunk;
if (Buffer.byteLength(raw) > 20_000) {
tooLarge = true;
res.writeHead(413, { "Content-Type": "application/json; charset=utf-8" });
res.end(JSON.stringify({ error: "payload_too_large" }));
}
});
req.on("end", () => {
if (tooLarge) return;
try {
const context = normalizeContext(JSON.parse(raw || "{}"));
const items = recommend(context);
res.writeHead(200, {
"Content-Type": "application/json; charset=utf-8",
"Cache-Control": "private, no-store"
});
res.end(JSON.stringify({ items }));
} catch {
res.writeHead(400, { "Content-Type": "application/json; charset=utf-8" });
res.end(JSON.stringify({ error: "invalid_json" }));
}
});
});
server.listen(3000, "127.0.0.1", () => {
console.log("Listening on http://127.0.0.1:3000");
});
启动并测试:
node server.js
curl -s http://127.0.0.1:3000/recommend \
-H 'content-type: application/json' \
-d '{
"locale": "zh-CN",
"period": "morning",
"recentCategories": ["developer"],
"dismissed": ["night-mode"]
}'
这个实现刻意保持简单。生产环境可以把规则替换成轻量模型,但输入边界不应随模型复杂度无限扩张。模型需要的每个特征都应有保留期限、用途说明和降级方案。
上线时最容易遗漏的边界
1. 个性化结果不能充当权限判断
客户端上下文是不可信输入。用户可以把 locale、类别或时间段改成任意值,因此这些字段只能影响推荐顺序、文案或布局,不能决定:
- 用户能否读取某项私有资源;
- 是否获得某个价格或金融额度;
- 是否绕过年龄、地区或合规限制;
- 是否拥有管理权限。
授权必须由经过验证的身份和服务端策略完成。无身份个性化不是身份认证的替代品。
2. 基础设施仍可能留下身份线索
即使应用代码不记录用户 ID,反向代理、CDN、WAF 和访问日志仍可能保存 IP、User-Agent 或完整请求正文。上线前应逐层检查:
- 禁止在日志中记录个性化请求正文;
- 缩短原始访问日志的保留时间;
- 对 IP 做截断、散列或边缘聚合,但要评估散列值是否仍可追踪;
- 不把上下文字段自动转发给第三方分析服务;
- 为调试采样设置明确比例和到期时间。
3. 指标不必回到个人层面
推荐系统仍然需要评估,但不一定需要建立个人点击历史。可以统计按小时或按粗粒度场景聚合的曝光、点击和隐藏次数,并设置最小样本阈值。
例如,与其保存“某设备依次点击了 A、B、C”,不如记录“上午时段的开发者内容获得 120 次曝光和 18 次点击”。这种指标无法完成精细归因,却足以发现排序规则是否完全失效。
同时要警惕反馈循环:如果系统只展示当前得分最高的类别,它将越来越确信用户只喜欢该类别。可以预留少量探索流量,或者允许用户直接重置偏好。
采用前的检查清单
无身份个性化适合内容排序、主题设置、搜索建议、入门路径和会话内辅助,但它会牺牲跨设备连续性与长期历史建模能力。采用前可以逐项确认:
- [ ] 每个输入字段都能解释它如何改变当前响应;
- [ ] 请求中没有账号、广告 ID、持久 Cookie 或设备指纹;
- [ ] 精确位置、时间和设备参数已经降到必要粒度;
- [ ] 原始上下文不会进入长期日志或第三方分析系统;
- [ ] 客户端字段只影响体验,不参与授权或高风险决策;
- [ ] 用户可以查看、重置或关闭本地偏好;
- [ ] 聚合指标设置了最小样本量和保留期限;
- [ ] 没有上下文时,系统仍能提供稳定、可用的默认结果。
真正有价值的目标不是把身份藏起来,而是让产品在根本不需要身份的情况下仍然有用。先从无状态规则、显式偏好和最少字段开始;只有在收益能够被验证时,才增加新的上下文信号。