不识别用户,也能做个性化:无账号、无持久 ID 的推荐设计

2026-09-30 30 预计阅读时间: 1 分钟
来源: medium.com AI 摘要 Original link

Disclaimer: This article is an AI-assisted summary. Read it together with the original source when precision matters. The summary may omit context, version differences, or edge cases and is not official documentation.

预计阅读时间:11 分钟

个性化并不必然意味着“先认出这个人”。即使没有账号、用户画像或跨会话追踪 ID,系统仍然可以根据当前请求、用户主动选择的偏好以及本次会话中的行为,调整内容排序和界面呈现。

在没有更多来源实现细节的前提下,下面给出一种可落地的设计:客户端提交有限的即时上下文,服务端执行无状态打分,用完即弃,并明确禁止将这些信号用于权限和价格判断。

把“个性化”和“识别身份”拆开

传统个性化系统经常从一个稳定标识开始:账号 ID、广告 ID、设备 ID,或者长期 Cookie。随后,系统围绕这个标识积累浏览、购买和点击历史。

无身份个性化采用不同的数据流。它不问“这个人过去是谁”,而只处理完成当前任务所需的信号,例如:

  • 用户在当前页面主动选择的语言、主题或内容类别;
  • 当前时间段、页面入口和设备能力;
  • 本次会话中刚刚浏览、隐藏或收藏的项目;
  • 保存在设备本地、由用户控制的界面偏好;
  • 不足以指向个人的粗粒度上下文,例如“移动端”而不是完整浏览器指纹。

这里的边界非常重要:匿名字段组合得足够多,仍可能形成指纹。IP 地址、精确位置、字体列表、屏幕尺寸、时区和行为时间戳叠加后,完全可能重新识别用户。因此,“没有 user_id 字段”不等于“没有身份风险”。

更稳妥的原则是只接受能够解释用途的字段:如果某个字段不能改变当前响应,就不要收集它。

一条无状态的数据路径

可以把请求设计成一次性的上下文计算:

客户端即时上下文
        ↓
字段校验与降精度
        ↓
无状态规则或模型打分
        ↓
返回排序结果
        ↓
丢弃原始上下文

客户端不需要发送账号或持久标识。例如,内容推荐接口可以只接收:

{
  "locale": "zh-CN",
  "period": "morning",
  "recentCategories": ["developer"],
  "dismissed": ["night-mode"]
}

这些信息描述的是“现在需要什么”,而不是“请求者是谁”。服务端也不必建立个人历史表。若要在页面切换间保留偏好,可以优先使用 sessionStorage;它的生命周期通常限于当前标签页会话。使用 localStorage 也不等同于建立账号,但会延长设备侧数据寿命,需要提供清除入口并说明用途。

可运行的 Node.js 无状态推荐接口

下面是一个只依赖 Node.js 标准库的最小示例。它不创建 Cookie、不接收用户 ID,也不保存请求正文。推荐结果只由本次提交的上下文计算。

将代码保存为 server.js,使用 Node.js 18 或更高版本运行:

const http = require("node:http");

const catalog = [
  {
    id: "docs-zh",
    title: "中文开发文档",
    categories: ["developer"],
    periods: ["morning", "afternoon"],
    locales: ["zh-CN"]
  },
  {
    id: "api-playground",
    title: "API 在线实验室",
    categories: ["developer", "api"],
    periods: ["afternoon", "evening"],
    locales: ["*"]
  },
  {
    id: "night-mode",
    title: "夜间阅读模式",
    categories: ["productivity"],
    periods: ["evening", "night"],
    locales: ["*"]
  },
  {
    id: "getting-started",
    title: "快速入门",
    categories: ["beginner"],
    periods: ["morning", "afternoon"],
    locales: ["*"]
  }
];

function normalizeContext(input) {
  const allowedPeriods = new Set(["morning", "afternoon", "evening", "night"]);

  return {
    locale: typeof input.locale === "string"
      ? input.locale.slice(0, 16)
      : "unknown",
    period: allowedPeriods.has(input.period)
      ? input.period
      : "unknown",
    recentCategories: Array.isArray(input.recentCategories)
      ? input.recentCategories
          .filter(value => typeof value === "string")
          .slice(0, 5)
      : [],
    dismissed: Array.isArray(input.dismissed)
      ? input.dismissed
          .filter(value => typeof value === "string")
          .slice(0, 20)
      : []
  };
}

function recommend(context) {
  const recent = new Set(context.recentCategories);
  const dismissed = new Set(context.dismissed);

  return catalog
    .filter(item => !dismissed.has(item.id))
    .map(item => {
      let score = 0;

      if (item.locales.includes(context.locale)) score += 3;
      if (item.locales.includes("*")) score += 1;
      if (item.categories.some(category => recent.has(category))) score += 2;
      if (item.periods.includes(context.period)) score += 1;

      return {
        id: item.id,
        title: item.title,
        score
      };
    })
    .sort((a, b) => b.score - a.score || a.id.localeCompare(b.id))
    .slice(0, 3);
}

const server = http.createServer((req, res) => {
  if (req.method !== "POST" || req.url !== "/recommend") {
    res.writeHead(404, { "Content-Type": "application/json; charset=utf-8" });
    return res.end(JSON.stringify({ error: "not_found" }));
  }

  let raw = "";
  let tooLarge = false;

  req.on("data", chunk => {
    if (tooLarge) return;
    raw += chunk;

    if (Buffer.byteLength(raw) > 20_000) {
      tooLarge = true;
      res.writeHead(413, { "Content-Type": "application/json; charset=utf-8" });
      res.end(JSON.stringify({ error: "payload_too_large" }));
    }
  });

  req.on("end", () => {
    if (tooLarge) return;

    try {
      const context = normalizeContext(JSON.parse(raw || "{}"));
      const items = recommend(context);

      res.writeHead(200, {
        "Content-Type": "application/json; charset=utf-8",
        "Cache-Control": "private, no-store"
      });
      res.end(JSON.stringify({ items }));
    } catch {
      res.writeHead(400, { "Content-Type": "application/json; charset=utf-8" });
      res.end(JSON.stringify({ error: "invalid_json" }));
    }
  });
});

server.listen(3000, "127.0.0.1", () => {
  console.log("Listening on http://127.0.0.1:3000");
});

启动并测试:

node server.js

curl -s http://127.0.0.1:3000/recommend \
  -H 'content-type: application/json' \
  -d '{
    "locale": "zh-CN",
    "period": "morning",
    "recentCategories": ["developer"],
    "dismissed": ["night-mode"]
  }'

这个实现刻意保持简单。生产环境可以把规则替换成轻量模型,但输入边界不应随模型复杂度无限扩张。模型需要的每个特征都应有保留期限、用途说明和降级方案。

上线时最容易遗漏的边界

1. 个性化结果不能充当权限判断

客户端上下文是不可信输入。用户可以把 locale、类别或时间段改成任意值,因此这些字段只能影响推荐顺序、文案或布局,不能决定:

  • 用户能否读取某项私有资源;
  • 是否获得某个价格或金融额度;
  • 是否绕过年龄、地区或合规限制;
  • 是否拥有管理权限。

授权必须由经过验证的身份和服务端策略完成。无身份个性化不是身份认证的替代品。

2. 基础设施仍可能留下身份线索

即使应用代码不记录用户 ID,反向代理、CDN、WAF 和访问日志仍可能保存 IP、User-Agent 或完整请求正文。上线前应逐层检查:

  • 禁止在日志中记录个性化请求正文;
  • 缩短原始访问日志的保留时间;
  • 对 IP 做截断、散列或边缘聚合,但要评估散列值是否仍可追踪;
  • 不把上下文字段自动转发给第三方分析服务;
  • 为调试采样设置明确比例和到期时间。

3. 指标不必回到个人层面

推荐系统仍然需要评估,但不一定需要建立个人点击历史。可以统计按小时或按粗粒度场景聚合的曝光、点击和隐藏次数,并设置最小样本阈值。

例如,与其保存“某设备依次点击了 A、B、C”,不如记录“上午时段的开发者内容获得 120 次曝光和 18 次点击”。这种指标无法完成精细归因,却足以发现排序规则是否完全失效。

同时要警惕反馈循环:如果系统只展示当前得分最高的类别,它将越来越确信用户只喜欢该类别。可以预留少量探索流量,或者允许用户直接重置偏好。

采用前的检查清单

无身份个性化适合内容排序、主题设置、搜索建议、入门路径和会话内辅助,但它会牺牲跨设备连续性与长期历史建模能力。采用前可以逐项确认:

  • [ ] 每个输入字段都能解释它如何改变当前响应;
  • [ ] 请求中没有账号、广告 ID、持久 Cookie 或设备指纹;
  • [ ] 精确位置、时间和设备参数已经降到必要粒度;
  • [ ] 原始上下文不会进入长期日志或第三方分析系统;
  • [ ] 客户端字段只影响体验,不参与授权或高风险决策;
  • [ ] 用户可以查看、重置或关闭本地偏好;
  • [ ] 聚合指标设置了最小样本量和保留期限;
  • [ ] 没有上下文时,系统仍能提供稳定、可用的默认结果。

真正有价值的目标不是把身份藏起来,而是让产品在根本不需要身份的情况下仍然有用。先从无状态规则、显式偏好和最少字段开始;只有在收益能够被验证时,才增加新的上下文信号。


相关推荐