Skip to content

SDD-002 · Evidence, các Model và các Engine ​

Một phần của SDD-002.

1. Evidence trước, kết luận sau ​

Evidence = fact gần quan sát ("sai 4/5 câu phân tích đa thức, >3 phút/câu, dùng hint 3 câu"), không lưu kết luận ("yếu đại số") — kết luận do Engine suy ra.

2. Learner Model (REQ-INT-02, REQ-LRN-01) ​

Computed state, không phải data warehouse (chi tiết growth control: SDD-007).

text
LearnerModel: Identity · Academic State · Capability State
· Learning Characteristics · Behavioral Patterns · Preferences · goals_summary (ref → Goal Model §16) · Confidence

Academic State theo Skill/KnowledgeNode (graph ở SDD-003): mỗi node {mastery: 0–1, confidence: 0–1, trajectory, last_evidence_at, evidence_count}. Mastery ≠ confidence — mastery 0.85 / confidence 0.20 = "có vẻ giỏi, chưa đủ bằng chứng"; UI phải phân biệt "chưa học" với "chưa đủ evidence" (REQ-VIS-04). Mastery cũng ≠ retention — mastery không tự giảm theo thời gian; trạng thái trí nhớ hiện tại là dimension riêng ở Retention Model (SDD-017, REQ-INT-23).

Learner Model = State + Evidence, không chỉ State (SRC-528, learner-v2). Một claim về learner phải trả lời được "vì sao tin": bản chụp mang theo hồ sơ bằng chứng — độ đa dạng (distinct_sources/distinct_types/by_type), mốc đầu–cuối (first_evidence_at/last_evidence_at) — chứ không chỉ total + avg_reliability. Freshness không lưu dạng số ngày: trường đó đổi theo đồng hồ chứ không theo learner, sẽ phá dedupe version; người đọc tự tính từ last_evidence_at. Thêm hai trục tách bạch: Coverage ≠ Mastery — mỗi môn mang nodes_total/coverage (mẫu số là cả graph, kể cả node chưa chạm) để avg_mastery cao trên coverage 5% không bị đọc thành "giỏi môn này"; và Behavior — tín hiệu cách-học quan sát được (sessions_abandoned, completion_rate, avg_response_ms) từ assessment_sessions/assessment_responses, chỉ đổi khi có hành vi mới. Còn thiếu có chủ đích: Transfer (dùng được ở context mới chưa) và Preferences quan sát được (help-seeking, retry pattern) — chưa có nguồn dữ liệu đáng tin, không bịa từ proxy yếu.

3. Learner Context Model ​

Trạng thái NGẮN HẠN — trả lời "điều gì đang xảy ra với learner ngay lúc này": academic (sắp thi), health (ốm/hồi phục), emotional (stress Hình học), time (tuần này 30'/ngày), family/environment. Là PROJECTION (materialized current-view) từ Context Event Registry (§5b), không phải source of truth (SRC-032): bảng learner_context_models chỉ là view hiện hành; sửa context = ghi event mới, không update tại chỗ. Mỗi datum bắt buộc mang {source, confidence} (§23 SRC-032 — "Geometry stress HIGH / learner self-report / 0.90" khác "MEDIUM / system inference / 0.52"). Goal + deadline KHÔNG còn là dữ liệu gốc ở đây — chỉ giữ active_goal_ref → Goal Model (§16). Available time/energy giữ tạm ở đây cho tới khi Constraint Model (Phase 2, Q-096) nhận — cấm xây feature mới trên 2 cột này. Cùng Learner Model nhưng context khác → recommendation khác (18 tháng vs 21 ngày trước kỳ thi). Context có TTL/decay (SDD-007 §12) + lifecycle rời rạc ACTIVE / RESOLVED / EXPIRED per event (§5b).

4. Parent Model (REQ-PAR-05) ​

ParentModel: Understanding of Learner · Education Beliefs · Expectations · Communication/Intervention Patterns · Availability · Constraints · Support Capacity · Goals. Lưu evidence + inferred tendency với confidence (intervention_style: high-direct-intervention, 0.68) — không label phán xét. Nguồn bổ sung phase sau: adaptive questions kiểu Member Intelligence sutucon (REQ-PAR-08 ⏳).

5. Evidence Registry (REQ-INT-01, REQ-MEN-03) ​

text
Evidence: evidence_id · learner_id · source · type · target(skill/node) · value
· timestamp · reliability · context · metadata · idempotency event_id

Reliability priors: standardized_assessment .95 · teacher .85 · learning_task .80 · mentor_observation .75 · parent_observation .60 · self-report .50 (điều chỉnh theo lịch sử). Immutable, append-only; correction qua evidence mới supersede. Idempotency key kiểu sutucon: product:<id>:final:<attempt>.

5b. Context Event Registry (REQ-INT-15, SRC-032) ​

Registry song song và cùng nguyên tắc với Evidence Registry (§5): immutable, append-only, idempotency theo event_id, sửa sai bằng event mới supersedes — không bao giờ sửa event cũ (lịch thi đổi = event EXAM_RESCHEDULED {old, new, supersedes: event_123}).

text
ContextEvent: event_id · learner_id · type (ILLNESS | EMOTIONAL_STATE | EXAM_SCHEDULED |
  EXAM_RESCHEDULED | EXAM_CANCELLED | TIME_AVAILABILITY | FAMILY_EVENT | SCHOOL_EVENT …)
· payload (domain/area/state/…) · severity · effective_from · effective_until?
· source (parent | learner_self | mentor | school | system_inference) · confidence
· supersedes? · lifecycle (ACTIVE | RESOLVED | EXPIRED) · observed_at
  • Không context nào tồn tại vô hạn: EXPIRED tự động theo effective_until/TTL type-default (SDD-007 §12); RESOLVED cần source xác nhận ("đã khỏi ốm").
  • Evidence mô tả những gì learner thể hiện; Context Event mô tả những gì xảy ra quanh learner. Learner ốm không làm mastery giảm — chỉ Recommendation đổi (SRC-032 §8).
  • Nguồn nạp: declarations tự nhiên của Nemo/Marlin (REQ-LRN-13, REQ-PAR-14 — interpretation qua AI Gateway, confirm trước khi ghi, không form 20 fields), school schedule, system inference. semester_exam_schedule (SDD-011 §6) trở thành projection từ EXAM_SCHEDULED/RESCHEDULED events.

6. Engines ​

text
Evidence Engine → Learner Engine → Assessment Engine → Recommendation Engine → Learning Engine
                + Context Engine · Goal Engine (§16) · Readiness Engine (§17)
                + Retention Engine (SDD-017) · Parent Engine · Progress Engine · Risk Engine · Depth Engine

Map thuật ngữ SRC-032 ↔ hệ hiện có: Learner Engine = vòng WF-04 (mastery.ts); Context Engine = build projection §3 từ registry §5b; Goal/Readiness Engine = §16/§17 (mới); Planning + Constraint Engine = Phase 2, đặt dưới Global Orchestrator (SDD-007 §9 — Q-096/Q-097); Assessment = §9 (information acquisition, đã live — phasing SRC-032 chỉ áp cho phần nâng cấp, Q-095); Learning = §11 (đã live, Turtle labs). Evidence/Parent/Progress/Risk engines ngoài scope SRC-032 — không bị supersede.

7. Learner Engine — mastery update (REQ-INT-03) ​

v1 ĐANG CHẠY (modules/knowledge/mastery.ts, tham số chuẩn ở reference/engines.md) là tập con của công thức đích dưới đây — Audit #013 (T-1) bắt được SDD mô tả công thức đích như thể đang chạy. Cụ thể v1: surprise = correct ? difficulty : 1 − difficulty; lr = (0.35 + 0.35·surprise) · 1/(1 + 0.6·evidence_count) · reliability; mastery' = mastery + lr·(outcome − mastery) với outcome nhị phân. Chưa có recency_weight và independence_weight; difficulty đi vào qua surprise chứ không phải trọng số riêng. Một câu đúng đầu tiên ở difficulty 0.5 đưa 0.60 → 0.81 (không phải 0.69 như ví dụ đích).

Công thức đích (Phase sau, chưa thi hành):

text
evidence_strength = reliability × recency_weight × difficulty_weight × independence_weight
α = evidence_strength × learning_rate
new_mastery = old_mastery × (1-α) + observed_performance × α

Ví dụ: 0.60 → (perf 0.90, α 0.30) → 0.69. Một câu đúng không được nhảy 0.4→0.9.

8. Confidence (REQ-INT-03) ​

v1 ĐANG CHẠY: confidence = 1 − 0.55^evidence_count (evidence_count chỉ tăng khi reliability ≥ 0.15). Không có source_diversity_factor, coverage_factor, consistency_factor: 10 câu đúng cùng dạng cho confidence 0.9975, 2 câu cho 0.70 — tức lời hứa chống false certainty ở đoạn dưới chưa tồn tại, và hai ngưỡng learner-facing đọc con số này (stateFor < 0.25 → unknown; GATE_CONFIDENCE_MIN = 0.4 ở entryGate.ts) gắn "solid" sau 3–4 câu cùng dạng. Việc kế (chưa có SRC): hiện thực source_diversity_factor từ distinct_types (§2, learner-v2) kèm test bất biến "10 câu cùng dạng → confidence ≤ 0.6".

Công thức đích:

text
confidence = (1 - exp(-weighted_evidence_count))
           × source_diversity_factor × coverage_factor × consistency_factor

10 câu đúng cùng dạng → mastery 0.85 / confidence 0.48; đa dạng context → 0.88. Chống false certainty.

9. Assessment Engine = information acquisition (REQ-INT-04) ​

Không "chấm bài" — trả lời "hệ thống còn cần biết gì?".

text
AssessmentValue = InformationGain × GoalRelevance × SkillImportance × DiagnosticPower ÷ Cost

Cost gồm thời gian, cognitive load, fatigue. Dừng khi confidence ≥ required_confidence — không ép bài test 100 câu cố định. Diagnosis engine phải có QA riêng (kế thừa 11 bảng diagnosis QA chuyenchon — FEAT-004): registry câu hỏi, review, validation policy.

10. Recommendation Engine (REQ-INT-05, REQ-LRN-02) ​

Candidates: Learn X · Review Y · Diagnostic Z · Skip · Practice · Ask Mentor · Rest · Prepare Exam.

text
Priority = ExpectedImpact × GoalRelevance × Urgency × PrerequisiteImportance
         × ProbabilityOfSuccess ÷ Cost   (+ ConfidenceNeed khi thiếu data)

Ưu tiên fix prerequisite gap đang block nhiều node (fractions 0.91 > new chapter 0.43). Mọi recommendation lưu reason + input_model_version + algorithm_version (§15). School engines chỉ tạo candidate — quyết định cuối ở Global Orchestrator (SDD-007 §9).

11. Learning Engine — HOW (REQ-INT-06, REQ-LRN-03/04) ​

Output: Explanation · Example · Scaffold · Practice · Worked Example · Hint · Challenge · Reflection. Difficulty Controller:

text
>90% đúng → tăng độ khó · 65–90% → tiếp tục · 40–65% → scaffold · <40% → prerequisite check

Scaffold vẫn fail → gọi Assessment Engine tìm prerequisite gap; cấm ném thêm 20 bài cùng dạng. Mode selection theo mastery/confidence kế thừa sutucon (guided/supported/practice/challenge — FEAT-009).

12. Depth Engine & Stopping Criteria (SRC-011; REQ-INT-11, REQ-LRN-07, REQ-SCH-03, REQ-PAR-06) ​

Depth levels: L0 Exposure · L1 Recognition · L2 Basic Application · L3 Independent Application · L4 Transfer · L5 Deep Mastery.

  • Mỗi learning objective có {required_depth, desired_mastery, required_confidence} — Turtle G7: L3/0.80/0.80; Shark thi chuyên: L4-L5/0.90/0.90. "Học kỹ" = target rõ ràng, không phải cảm giác. (Hiện thân legacy: math_target_blueprint conditional/specialized — FEAT-038.)
  • STOP khi mastery ≥ threshold AND confidence ≥ threshold AND depth ≥ required AND evidence diversity ≥ min.
  • Diminishing returns: khi ExpectedLearningGain / TimeCost < threshold → chuyển sang node impact cao hơn. Không tối đa hóa từng node — tối đa hóa tổng thể dưới giới hạn thời gian.
  • Overlearning protection: phát hiện luyện lặp skill đã ổn định / extension không phục vụ goal / đào sâu khi còn critical gap khác.
  • Depth Budget theo goal: critical L4/L5 · important L3/L4 · peripheral L2/L3.
  • Parent-facing: giải thích target/current/mastery/confidence + evidence; muốn nâng L3→L4 phải thấy cost (4–6h) và trade-off với Skill B (REQ-PAR-06). Câu hỏi trung tâm: "học sâu thêm ở đây có còn là next best use of learner time?"

13. Parent Engine (REQ-PAR-01/02/04) ​

text
parent_action = impact × feasibility × parent_capacity ÷ intervention_cost

Ví dụ: learner fail lặp → Learner Engine tìm ra prerequisite gap → Parent Engine biết parent hay "thêm lớp" → recommend "Chưa cần thêm lớp; 3 ngày xử lý prerequisite X" — chống intervention error, output "không nên làm gì" là hợp lệ.

Định tuyến observation (SRC-032, REQ-PAR-14): mỗi parent/mentor observation được phân loại → {skill evidence | context event | cả hai} — "con sợ hình" = EMOTIONAL_STATE (§5b) + tín hiệu affect về Geometry; ngừng dual-write node-rỗng vào Evidence Registry (Q-101).

14. Main Loop & Workflow (REQ-INT-13) ​

text
Evidence → Registry → Learner Engine update → đủ confidence?
  ├─ No → Assessment Engine (thu evidence)
  └─ Yes → Recommendation → Learning Engine → Experience → New Evidence ↺

AssessmentCompleted → Workflow: validate → persist → update skills → propagate prerequisite impact → recompute confidence → detect changes → recommendations → update context → notify. Bước độc lập fan-out qua Queue; orchestration qua Workflow. Không chạy trong 1 HTTP request.

Model Update Policy (REQ-INT-18, SRC-032 §27-28) — không phải event nào cũng update mọi model; nguyên tắc thứ tự: update Context/Goal/Readiness TRƯỚC, regenerate recommendation CUỐI:

EventModels updateEngines re-runKhông đụng
AssessmentCompleted / practice sai-đúngLearner, Retention (SDD-017), Readiness (targets liên quan)Learner → Retention → Readiness → (Planning) → RecommendationGoal, Context
RETRIEVAL_SUCCEEDED / RETRIEVAL_FAILED / review completedRetention (node + next_review_window)Retention → RecommendationLearner mastery lịch sử, Goal, Context
ContextEventRecorded: ILLNESSContext (+Constraint P2), PlanContext → Planning → RecommendationLearner mastery
ContextEventRecorded: EMOTIONAL_STATEContextContext → Recommendation (đổi intervention, không dừng môn)Learner, Goal
EXAM_RESCHEDULEDContext, Goal (deadline), Readiness urgency, PlanContext → Goal → Readiness → Planning → RecommendationLearner mastery
GoalChanged (khai mới/bỏ)Goal, Readiness, PlanGoal → Readiness → Planning → RecommendationLearner, Context

Queues thêm event types: ContextEventRecorded, GoalChanged, ConstraintChanged (SDD-007 §15).

15. Versioning & Explainability (REQ-INT-14) ​

Generalize cho cả 6 models (SRC-032 §24): ModelVersion {model_kind, version, generated_at, evidence_cutoff, algorithm_version, state}. Recommendation lưu {reason, input_model_version, goal_model_version, context_version, plan_version, algorithm_version} vào recommendation_log → recommendation diff: so 2 lần chạy, quy được về event/model-version gây thay đổi ("khác hôm qua vì: +lịch thi học kỳ, +ILLNESS event, −60' available, +Geometry readiness") — SRC-032 §24-25. AI output lưu {model, prompt_version, input_ref, output, confidence, evaluation_state} (QG-010).

16. Goal Model & Goal Engine (REQ-INT-16, SRC-032 §9-10) ​

Trạng thái: đã hiện thực — bảng learner_goal_entries, engine thuần modules/goals/engine.ts, chạy qua runGoalEngine() và ghi version vào Goal Model store (§19).

"Learner đang cố đạt điều gì?" — model độc lập, được phép rỗng (REQ-ONB-05, RISK-009 — không ép mục tiêu).

text
Goal: goal_id · learner_id · title · kind (exam_school | ielts | sat | ap | semester_exam
  | foundation_repair | custom) · parent_goal_id? (goal tree: "Đỗ CSP" → "Toán chuyên readiness"
  → "Geometry") · blueprint_id? (goal template — SDD-011 §2b) · priority · deadline?
  · dependencies[] · conflict_flags[] · source (parent | learner_self | enrollment | system)
  · status (active | achieved | dropped) · version
  • Ngày thi = deadline trên goal entry — nguồn duy nhất; thay đổi qua event EXAM_RESCHEDULED (§5b), không có bảng lịch thi thứ hai (hợp nhất 3 điểm nhập cũ: semester_exam_schedule, exam_date_approx, declarations).
  • Goal Engine xử lý: priority (operational vs strategic — thi học kỳ 5 ngày ưu tiên vận hành, CSP 180 ngày ưu tiên chiến lược, không nhầm "goal lớn nhất = chỉ học chuyên"), dependency, conflict (2 goals cạnh tranh thời gian → flag cho Planning/Orchestrator; parent-goal vs learner-goal khác nhau → escalate cho parent, hệ không tự quyết — Q-099).
  • Nguồn: learner_exam_targets (SDD-011 §2b — mỗi target sinh goal tree; môn không chuyên = attribute của parent goal, Q-102), enrollment, declarations, system inference.

17. Readiness Model & Readiness Engine (REQ-INT-17, REQ-LRN-05, SRC-032 §13-14) ​

"Còn cách Goal bao xa?" — luôn gắn một Target (exam target / semester exam / blueprint depth target §12). Readiness ≠ mastery: "Math khá" nhưng "CSP readiness thấp" là hợp lệ.

text
Readiness: learner_id · target_ref (goal_id + blueprint) · score 0–1 · confidence 0–1
· gap_map { node → (required, current, gap) } · largest_uncertainty[]
· next_assessment_priority[] · version
  • Input: Goal Model (§16) + depth targets (§12 {required_depth, desired_mastery, required_confidence}) + Academic State (§2) + assessment evidence.
  • Output không chỉ một con số: readiness 0.61 + primary gaps (Geometry, Number Theory) + largest uncertainty (Combinatorics — confidence thấp) + next_assessment_priority → Assessment Engine (§9) tiêu thụ priority này trong công thức AssessmentValue (Readiness sinh nhu cầu đo, Assessment chọn cách đo — biên giới rõ).
  • SDD-008 §5 Goal View = pure view đọc từ model này (thêm hiển thị uncertainty); cockpit computeReadiness đọc/ghi model, không tính ad-hoc.
  • Depth Engine §12 về taxonomy: target/gap side thuộc Readiness, allocation/stopping side thuộc Planning (Phase 2) — giữ §12 làm spec nội dung, không đổi code (Q-098).

18. Evidence Reliability & Rapid Guessing Prevention (REQ-INT-21/22, SRC-038) ​

Answer ≠ Evidence of Knowledge — hệ đánh giá cả đáp án lẫn độ tin cậy của hành vi trả lời, để learner chọn bừa không phá Learner Model (không tạo false low-mastery, không làm Recommendation bắt học lại thứ đã biết).

  • Detection: mỗi attempt tính rapid = response_ms < max(2.5s, 20% expected_time); theo dõi chuỗi rapid liên tiếp trong session, accuracy ~ mức random, hint usage, confidence khai. v1 đang chạy: RAPID_MS = 2500 cố định (mastery.ts); expected_time chưa tồn tại ở đâu trong code hay data (Audit #013, T-1).
  • Reliability per attempt (lưu cùng assessment_responses/evidence metadata): {correctness, response_ms, confidence?, guessing_probability, reliability_score}. Heuristic v1: bình thường = 1.0 · rapid đơn lẻ = 0.4 · chuỗi rapid ≥ 3 = 0.05 (gần như không ảnh hưởng mastery) · "Chưa biết" = 0.7 (tín hiệu trung thực, performance 0).
  • Mastery update (§7): lr *= reliability_score — đúng-nhanh-bừa là weak evidence; đúng + đủ thời gian + confidence cao là strong evidence.
  • Progressive intervention (tăng dần, không trừng phạt): (1) nhắc nhẹ "chậm lại, đọc kỹ đề" → (2) nút "Chưa biết" hợp lệ → (3) hỏi confidence → (4) delay submit ngắn → (5) đánh dấu assessment INSUFFICIENT_RELIABLE_EVIDENCE.
  • Assessment Engine: khi tỷ lệ rapid ≥ 50% session → kết luận INSUFFICIENT_RELIABLE_EVIDENCE thay vì "mastery 25%"; node giữ mastery UNKNOWN / confidence LOW; Recommendation ưu tiên thu evidence mới (đo lại lúc khác) chứ không bắt học lại toàn bộ. v1 đang chạy: diagnostic ≥ 50% rapid ghi evidence_reliability='insufficient' và hiện màn "chưa đủ tin cậy", nhưng knowledge/routes.ts vẫn cập nhật mastery với reliability = min(rel, 0.05) (bước rất nhỏ) thay vì giữ UNKNOWN — khác chữ "giữ" ở đây (Audit #013, T-1).
  • Metrics: rapid guess rate ↓, evidence reliability ↑, false low-mastery ↓, % assessment có dữ liệu đáng tin ↑.