Rules first
The game engine enumerates legal Dou Dizhu actions before AI scoring. The policy never invents an illegal play.
The web game and iOS app now load the same shared Dou Dizhu policy artifact. Normal play still runs on-device or in-browser, and both clients report settlement experience into one central learning pool.
The AI model should not require a web server for every decision. Web and iOS fetch the active shared policy, then score legal moves locally so solo play stays fast and inexpensive. At the end of a round, each client sends a compact experience event to the same backend pool. Future training runs can use that combined web plus iOS history to publish the next active policy.
The game engine enumerates legal Dou Dizhu actions before AI scoring. The policy never invents an illegal play.
The model scores legal candidates using hand pressure, partner state, landlord danger, bombs, rockets, and immediate-win value.
A farmer can beat another farmer when it helps the team, such as finishing the hand or blocking a dangerous landlord. Wasteful partner cuts are penalized.
Round outcomes from web and iOS are stored together by model id, platform, roles, outcome, decision count, and compact settlement summary.
The current linear model can later be replaced by a deeper self-play model or Core ML artifact without changing the legal-action boundary.
This is the first transparent self-play artifact, designed to be easy to inspect and safe to ship before a heavier neural model.
| Model id | ddz_policy_v1_team3 |
|---|---|
| Model type | linear-self-play-v1 |
| Training method | evolutionary-self-play-linear-policy-hard-guardrails |
| Training run | 80 episodes, 6 generations, seed 2599, best score 2.6195 |
| Primary features | immediate_win, preservation_cost, landlord_danger, farmer_blocks_landlord, farmer_helps_partner_finish, farmer_unneeded_partner_cut |
| References | RLCard for legal-action environment framing, and DouZero for self-play Dou Dizhu direction. |
| Limitation | This is not a full DouZero neural policy. It is the first inspectable model path and deployment contract. |
Expand self-play episodes with fuller round outcomes, role rewards, and better opponent sampling.
Train candidates from the shared experience pool, run policy-vs-policy matches, and publish win-rate summaries on this page.
Evaluate a DouZero-style learned policy and export it for mobile inference when quality justifies the added complexity.
Jest can be added for JavaScript UI/model-card checks. Native gameplay remains covered by XCTest and Python trainer tests.