SDD-002 · Evidence, các Model và các Engine
Một phần của SDD-002.
1. Evidence trước, kết luận sau
Evidence = fact gần quan sát ("sai 4/5 câu phân tích đa thức, >3 phút/câu, dùng hint 3 câu"), không lưu kết luận ("yếu đại số") — kết luận do Engine suy ra.
2. Learner Model (REQ-INT-02, REQ-LRN-01)
Computed state, không phải data warehouse (chi tiết growth control: SDD-007).
LearnerModel: Identity · Academic State · Capability State
· Learning Characteristics · Behavioral Patterns · Preferences · goals_summary (ref → Goal Model §16) · ConfidenceAcademic State theo Skill/KnowledgeNode (graph ở SDD-003): mỗi node {mastery: 0–1, confidence: 0–1, trajectory, last_evidence_at, evidence_count}. Mastery ≠ confidence — mastery 0.85 / confidence 0.20 = "có vẻ giỏi, chưa đủ bằng chứng"; UI phải phân biệt "chưa học" với "chưa đủ evidence" (REQ-VIS-04). Mastery cũng ≠ retention — mastery không tự giảm theo thời gian; trạng thái trí nhớ hiện tại là dimension riêng ở Retention Model (SDD-017, REQ-INT-23).
Learner Model = State + Evidence, không chỉ State (SRC-528, learner-v2). Một claim về learner phải trả lời được "vì sao tin": bản chụp mang theo hồ sơ bằng chứng — độ đa dạng (distinct_sources/distinct_types/by_type), mốc đầu–cuối (first_evidence_at/last_evidence_at) — chứ không chỉ total + avg_reliability. Freshness không lưu dạng số ngày: trường đó đổi theo đồng hồ chứ không theo learner, sẽ phá dedupe version; người đọc tự tính từ last_evidence_at. Thêm hai trục tách bạch: Coverage ≠ Mastery — mỗi môn mang nodes_total/coverage (mẫu số là cả graph, kể cả node chưa chạm) để avg_mastery cao trên coverage 5% không bị đọc thành "giỏi môn này"; và Behavior — tín hiệu cách-học quan sát được (sessions_abandoned, completion_rate, avg_response_ms) từ assessment_sessions/assessment_responses, chỉ đổi khi có hành vi mới. Còn thiếu có chủ đích: Transfer (dùng được ở context mới chưa) và Preferences quan sát được (help-seeking, retry pattern) — chưa có nguồn dữ liệu đáng tin, không bịa từ proxy yếu.
3. Learner Context Model
Trạng thái NGẮN HẠN — trả lời "điều gì đang xảy ra với learner ngay lúc này": academic (sắp thi), health (ốm/hồi phục), emotional (stress Hình học), time (tuần này 30'/ngày), family/environment. Là PROJECTION (materialized current-view) từ Context Event Registry (§5b), không phải source of truth (SRC-032): bảng learner_context_models chỉ là view hiện hành; sửa context = ghi event mới, không update tại chỗ. Mỗi datum bắt buộc mang {source, confidence} (§23 SRC-032 — "Geometry stress HIGH / learner self-report / 0.90" khác "MEDIUM / system inference / 0.52"). Goal + deadline KHÔNG còn là dữ liệu gốc ở đây — chỉ giữ active_goal_ref → Goal Model (§16). Available time/energy giữ tạm ở đây cho tới khi Constraint Model (Phase 2, Q-096) nhận — cấm xây feature mới trên 2 cột này. Cùng Learner Model nhưng context khác → recommendation khác (18 tháng vs 21 ngày trước kỳ thi). Context có TTL/decay (SDD-007 §12) + lifecycle rời rạc ACTIVE / RESOLVED / EXPIRED per event (§5b).
4. Parent Model (REQ-PAR-05)
ParentModel: Understanding of Learner · Education Beliefs · Expectations · Communication/Intervention Patterns · Availability · Constraints · Support Capacity · Goals. Lưu evidence + inferred tendency với confidence (intervention_style: high-direct-intervention, 0.68) — không label phán xét. Nguồn bổ sung phase sau: adaptive questions kiểu Member Intelligence sutucon (REQ-PAR-08 ⏳).
5. Evidence Registry (REQ-INT-01, REQ-MEN-03)
Evidence: evidence_id · learner_id · source · type · target(skill/node) · value
· timestamp · reliability · context · metadata · idempotency event_idReliability priors: standardized_assessment .95 · teacher .85 · learning_task .80 · mentor_observation .75 · parent_observation .60 · self-report .50 (điều chỉnh theo lịch sử). Immutable, append-only; correction qua evidence mới supersede. Idempotency key kiểu sutucon: product:<id>:final:<attempt>.
5b. Context Event Registry (REQ-INT-15, SRC-032)
Registry song song và cùng nguyên tắc với Evidence Registry (§5): immutable, append-only, idempotency theo event_id, sửa sai bằng event mới supersedes — không bao giờ sửa event cũ (lịch thi đổi = event EXAM_RESCHEDULED {old, new, supersedes: event_123}).
ContextEvent: event_id · learner_id · type (ILLNESS | EMOTIONAL_STATE | EXAM_SCHEDULED |
EXAM_RESCHEDULED | EXAM_CANCELLED | TIME_AVAILABILITY | FAMILY_EVENT | SCHOOL_EVENT …)
· payload (domain/area/state/…) · severity · effective_from · effective_until?
· source (parent | learner_self | mentor | school | system_inference) · confidence
· supersedes? · lifecycle (ACTIVE | RESOLVED | EXPIRED) · observed_at- Không context nào tồn tại vô hạn:
EXPIREDtự động theoeffective_until/TTL type-default (SDD-007 §12);RESOLVEDcần source xác nhận ("đã khỏi ốm"). - Evidence mô tả những gì learner thể hiện; Context Event mô tả những gì xảy ra quanh learner. Learner ốm không làm mastery giảm — chỉ Recommendation đổi (SRC-032 §8).
- Nguồn nạp: declarations tự nhiên của Nemo/Marlin (REQ-LRN-13, REQ-PAR-14 — interpretation qua AI Gateway, confirm trước khi ghi, không form 20 fields), school schedule, system inference.
semester_exam_schedule(SDD-011 §6) trở thành projection từ EXAM_SCHEDULED/RESCHEDULED events.
6. Engines
Evidence Engine → Learner Engine → Assessment Engine → Recommendation Engine → Learning Engine
+ Context Engine · Goal Engine (§16) · Readiness Engine (§17)
+ Retention Engine (SDD-017) · Parent Engine · Progress Engine · Risk Engine · Depth EngineMap thuật ngữ SRC-032 ↔ hệ hiện có: Learner Engine = vòng WF-04 (mastery.ts); Context Engine = build projection §3 từ registry §5b; Goal/Readiness Engine = §16/§17 (mới); Planning + Constraint Engine = Phase 2, đặt dưới Global Orchestrator (SDD-007 §9 — Q-096/Q-097); Assessment = §9 (information acquisition, đã live — phasing SRC-032 chỉ áp cho phần nâng cấp, Q-095); Learning = §11 (đã live, Turtle labs). Evidence/Parent/Progress/Risk engines ngoài scope SRC-032 — không bị supersede.
7. Learner Engine — mastery update (REQ-INT-03)
v1 ĐANG CHẠY (
modules/knowledge/mastery.ts, tham số chuẩn ở reference/engines.md) là tập con của công thức đích dưới đây — Audit #013 (T-1) bắt được SDD mô tả công thức đích như thể đang chạy. Cụ thể v1:surprise = correct ? difficulty : 1 − difficulty;lr = (0.35 + 0.35·surprise) · 1/(1 + 0.6·evidence_count) · reliability;mastery' = mastery + lr·(outcome − mastery)với outcome nhị phân. Chưa córecency_weightvàindependence_weight;difficultyđi vào quasurprisechứ không phải trọng số riêng. Một câu đúng đầu tiên ở difficulty 0.5 đưa 0.60 → 0.81 (không phải 0.69 như ví dụ đích).
Công thức đích (Phase sau, chưa thi hành):
evidence_strength = reliability × recency_weight × difficulty_weight × independence_weight
α = evidence_strength × learning_rate
new_mastery = old_mastery × (1-α) + observed_performance × αVí dụ: 0.60 → (perf 0.90, α 0.30) → 0.69. Một câu đúng không được nhảy 0.4→0.9.
8. Confidence (REQ-INT-03)
v1 ĐANG CHẠY:
confidence = 1 − 0.55^evidence_count(evidence_countchỉ tăng khireliability ≥ 0.15). Không cósource_diversity_factor,coverage_factor,consistency_factor: 10 câu đúng cùng dạng cho confidence 0.9975, 2 câu cho 0.70 — tức lời hứa chống false certainty ở đoạn dưới chưa tồn tại, và hai ngưỡng learner-facing đọc con số này (stateFor< 0.25 → unknown;GATE_CONFIDENCE_MIN = 0.4ởentryGate.ts) gắn "solid" sau 3–4 câu cùng dạng. Việc kế (chưa có SRC): hiện thựcsource_diversity_factortừdistinct_types(§2, learner-v2) kèm test bất biến "10 câu cùng dạng → confidence ≤ 0.6".
Công thức đích:
confidence = (1 - exp(-weighted_evidence_count))
× source_diversity_factor × coverage_factor × consistency_factor10 câu đúng cùng dạng → mastery 0.85 / confidence 0.48; đa dạng context → 0.88. Chống false certainty.
9. Assessment Engine = information acquisition (REQ-INT-04)
Không "chấm bài" — trả lời "hệ thống còn cần biết gì?".
AssessmentValue = InformationGain × GoalRelevance × SkillImportance × DiagnosticPower ÷ CostCost gồm thời gian, cognitive load, fatigue. Dừng khi confidence ≥ required_confidence — không ép bài test 100 câu cố định. Diagnosis engine phải có QA riêng (kế thừa 11 bảng diagnosis QA chuyenchon — FEAT-004): registry câu hỏi, review, validation policy.
10. Recommendation Engine (REQ-INT-05, REQ-LRN-02)
Candidates: Learn X · Review Y · Diagnostic Z · Skip · Practice · Ask Mentor · Rest · Prepare Exam.
Priority = ExpectedImpact × GoalRelevance × Urgency × PrerequisiteImportance
× ProbabilityOfSuccess ÷ Cost (+ ConfidenceNeed khi thiếu data)Ưu tiên fix prerequisite gap đang block nhiều node (fractions 0.91 > new chapter 0.43). Mọi recommendation lưu reason + input_model_version + algorithm_version (§15). School engines chỉ tạo candidate — quyết định cuối ở Global Orchestrator (SDD-007 §9).
11. Learning Engine — HOW (REQ-INT-06, REQ-LRN-03/04)
Output: Explanation · Example · Scaffold · Practice · Worked Example · Hint · Challenge · Reflection. Difficulty Controller:
>90% đúng → tăng độ khó · 65–90% → tiếp tục · 40–65% → scaffold · <40% → prerequisite checkScaffold vẫn fail → gọi Assessment Engine tìm prerequisite gap; cấm ném thêm 20 bài cùng dạng. Mode selection theo mastery/confidence kế thừa sutucon (guided/supported/practice/challenge — FEAT-009).
12. Depth Engine & Stopping Criteria (SRC-011; REQ-INT-11, REQ-LRN-07, REQ-SCH-03, REQ-PAR-06)
Depth levels: L0 Exposure · L1 Recognition · L2 Basic Application · L3 Independent Application · L4 Transfer · L5 Deep Mastery.
- Mỗi learning objective có
{required_depth, desired_mastery, required_confidence}— Turtle G7: L3/0.80/0.80; Shark thi chuyên: L4-L5/0.90/0.90. "Học kỹ" = target rõ ràng, không phải cảm giác. (Hiện thân legacy:math_target_blueprintconditional/specialized — FEAT-038.) - STOP khi
mastery ≥ threshold AND confidence ≥ threshold AND depth ≥ required AND evidence diversity ≥ min. - Diminishing returns: khi
ExpectedLearningGain / TimeCost < threshold→ chuyển sang node impact cao hơn. Không tối đa hóa từng node — tối đa hóa tổng thể dưới giới hạn thời gian. - Overlearning protection: phát hiện luyện lặp skill đã ổn định / extension không phục vụ goal / đào sâu khi còn critical gap khác.
- Depth Budget theo goal: critical L4/L5 · important L3/L4 · peripheral L2/L3.
- Parent-facing: giải thích target/current/mastery/confidence + evidence; muốn nâng L3→L4 phải thấy cost (4–6h) và trade-off với Skill B (REQ-PAR-06). Câu hỏi trung tâm: "học sâu thêm ở đây có còn là next best use of learner time?"
13. Parent Engine (REQ-PAR-01/02/04)
parent_action = impact × feasibility × parent_capacity ÷ intervention_costVí dụ: learner fail lặp → Learner Engine tìm ra prerequisite gap → Parent Engine biết parent hay "thêm lớp" → recommend "Chưa cần thêm lớp; 3 ngày xử lý prerequisite X" — chống intervention error, output "không nên làm gì" là hợp lệ.
Định tuyến observation (SRC-032, REQ-PAR-14): mỗi parent/mentor observation được phân loại → {skill evidence | context event | cả hai} — "con sợ hình" = EMOTIONAL_STATE (§5b) + tín hiệu affect về Geometry; ngừng dual-write node-rỗng vào Evidence Registry (Q-101).
14. Main Loop & Workflow (REQ-INT-13)
Evidence → Registry → Learner Engine update → đủ confidence?
├─ No → Assessment Engine (thu evidence)
└─ Yes → Recommendation → Learning Engine → Experience → New Evidence ↺AssessmentCompleted → Workflow: validate → persist → update skills → propagate prerequisite impact → recompute confidence → detect changes → recommendations → update context → notify. Bước độc lập fan-out qua Queue; orchestration qua Workflow. Không chạy trong 1 HTTP request.
Model Update Policy (REQ-INT-18, SRC-032 §27-28) — không phải event nào cũng update mọi model; nguyên tắc thứ tự: update Context/Goal/Readiness TRƯỚC, regenerate recommendation CUỐI:
| Event | Models update | Engines re-run | Không đụng |
|---|---|---|---|
AssessmentCompleted / practice sai-đúng | Learner, Retention (SDD-017), Readiness (targets liên quan) | Learner → Retention → Readiness → (Planning) → Recommendation | Goal, Context |
RETRIEVAL_SUCCEEDED / RETRIEVAL_FAILED / review completed | Retention (node + next_review_window) | Retention → Recommendation | Learner mastery lịch sử, Goal, Context |
ContextEventRecorded: ILLNESS | Context (+Constraint P2), Plan | Context → Planning → Recommendation | Learner mastery |
ContextEventRecorded: EMOTIONAL_STATE | Context | Context → Recommendation (đổi intervention, không dừng môn) | Learner, Goal |
EXAM_RESCHEDULED | Context, Goal (deadline), Readiness urgency, Plan | Context → Goal → Readiness → Planning → Recommendation | Learner mastery |
GoalChanged (khai mới/bỏ) | Goal, Readiness, Plan | Goal → Readiness → Planning → Recommendation | Learner, Context |
Queues thêm event types: ContextEventRecorded, GoalChanged, ConstraintChanged (SDD-007 §15).
15. Versioning & Explainability (REQ-INT-14)
Generalize cho cả 6 models (SRC-032 §24): ModelVersion {model_kind, version, generated_at, evidence_cutoff, algorithm_version, state}. Recommendation lưu {reason, input_model_version, goal_model_version, context_version, plan_version, algorithm_version} vào recommendation_log → recommendation diff: so 2 lần chạy, quy được về event/model-version gây thay đổi ("khác hôm qua vì: +lịch thi học kỳ, +ILLNESS event, −60' available, +Geometry readiness") — SRC-032 §24-25. AI output lưu {model, prompt_version, input_ref, output, confidence, evaluation_state} (QG-010).
16. Goal Model & Goal Engine (REQ-INT-16, SRC-032 §9-10)
Trạng thái: đã hiện thực — bảng learner_goal_entries, engine thuần modules/goals/engine.ts, chạy qua runGoalEngine() và ghi version vào Goal Model store (§19).
"Learner đang cố đạt điều gì?" — model độc lập, được phép rỗng (REQ-ONB-05, RISK-009 — không ép mục tiêu).
Goal: goal_id · learner_id · title · kind (exam_school | ielts | sat | ap | semester_exam
| foundation_repair | custom) · parent_goal_id? (goal tree: "Đỗ CSP" → "Toán chuyên readiness"
→ "Geometry") · blueprint_id? (goal template — SDD-011 §2b) · priority · deadline?
· dependencies[] · conflict_flags[] · source (parent | learner_self | enrollment | system)
· status (active | achieved | dropped) · version- Ngày thi = deadline trên goal entry — nguồn duy nhất; thay đổi qua event
EXAM_RESCHEDULED(§5b), không có bảng lịch thi thứ hai (hợp nhất 3 điểm nhập cũ:semester_exam_schedule,exam_date_approx, declarations). - Goal Engine xử lý: priority (operational vs strategic — thi học kỳ 5 ngày ưu tiên vận hành, CSP 180 ngày ưu tiên chiến lược, không nhầm "goal lớn nhất = chỉ học chuyên"), dependency, conflict (2 goals cạnh tranh thời gian → flag cho Planning/Orchestrator; parent-goal vs learner-goal khác nhau → escalate cho parent, hệ không tự quyết — Q-099).
- Nguồn:
learner_exam_targets(SDD-011 §2b — mỗi target sinh goal tree; môn không chuyên = attribute của parent goal, Q-102), enrollment, declarations, system inference.
17. Readiness Model & Readiness Engine (REQ-INT-17, REQ-LRN-05, SRC-032 §13-14)
"Còn cách Goal bao xa?" — luôn gắn một Target (exam target / semester exam / blueprint depth target §12). Readiness ≠ mastery: "Math khá" nhưng "CSP readiness thấp" là hợp lệ.
Readiness: learner_id · target_ref (goal_id + blueprint) · score 0–1 · confidence 0–1
· gap_map { node → (required, current, gap) } · largest_uncertainty[]
· next_assessment_priority[] · version- Input: Goal Model (§16) + depth targets (§12
{required_depth, desired_mastery, required_confidence}) + Academic State (§2) + assessment evidence. - Output không chỉ một con số: readiness 0.61 + primary gaps (Geometry, Number Theory) + largest uncertainty (Combinatorics — confidence thấp) +
next_assessment_priority→ Assessment Engine (§9) tiêu thụ priority này trong công thức AssessmentValue (Readiness sinh nhu cầu đo, Assessment chọn cách đo — biên giới rõ). - SDD-008 §5 Goal View = pure view đọc từ model này (thêm hiển thị uncertainty); cockpit
computeReadinessđọc/ghi model, không tính ad-hoc. - Depth Engine §12 về taxonomy: target/gap side thuộc Readiness, allocation/stopping side thuộc Planning (Phase 2) — giữ §12 làm spec nội dung, không đổi code (Q-098).
18. Evidence Reliability & Rapid Guessing Prevention (REQ-INT-21/22, SRC-038)
Answer ≠ Evidence of Knowledge — hệ đánh giá cả đáp án lẫn độ tin cậy của hành vi trả lời, để learner chọn bừa không phá Learner Model (không tạo false low-mastery, không làm Recommendation bắt học lại thứ đã biết).
- Detection: mỗi attempt tính
rapid = response_ms < max(2.5s, 20% expected_time); theo dõi chuỗi rapid liên tiếp trong session, accuracy ~ mức random, hint usage, confidence khai. v1 đang chạy:RAPID_MS = 2500cố định (mastery.ts);expected_timechưa tồn tại ở đâu trong code hay data (Audit #013, T-1). - Reliability per attempt (lưu cùng assessment_responses/evidence metadata):
{correctness, response_ms, confidence?, guessing_probability, reliability_score}. Heuristic v1: bình thường = 1.0 · rapid đơn lẻ = 0.4 · chuỗi rapid ≥ 3 = 0.05 (gần như không ảnh hưởng mastery) · "Chưa biết" = 0.7 (tín hiệu trung thực, performance 0). - Mastery update (§7):
lr *= reliability_score— đúng-nhanh-bừa là weak evidence; đúng + đủ thời gian + confidence cao là strong evidence. - Progressive intervention (tăng dần, không trừng phạt): (1) nhắc nhẹ "chậm lại, đọc kỹ đề" → (2) nút "Chưa biết" hợp lệ → (3) hỏi confidence → (4) delay submit ngắn → (5) đánh dấu assessment
INSUFFICIENT_RELIABLE_EVIDENCE. - Assessment Engine: khi tỷ lệ rapid ≥ 50% session → kết luận
INSUFFICIENT_RELIABLE_EVIDENCEthay vì "mastery 25%"; node giữmastery UNKNOWN / confidence LOW; Recommendation ưu tiên thu evidence mới (đo lại lúc khác) chứ không bắt học lại toàn bộ. v1 đang chạy: diagnostic ≥ 50% rapid ghievidence_reliability='insufficient'và hiện màn "chưa đủ tin cậy", nhưngknowledge/routes.tsvẫn cập nhật mastery vớireliability = min(rel, 0.05)(bước rất nhỏ) thay vì giữ UNKNOWN — khác chữ "giữ" ở đây (Audit #013, T-1). - Metrics: rapid guess rate ↓, evidence reliability ↑, false low-mastery ↓, % assessment có dữ liệu đáng tin ↑.