AImethodologyfor free DouDizhu.

The web game and iOS app now load the same shared Dou Dizhu policy artifact. Normal play still runs on-device or in-browser, and both clients report settlement experience into one central learning pool.

Dou Dizhu card game banner
Runtime Web and iOS score legal bot moves with the same policy schema and active model id.
Hosting Cloudflare Pages Functions serve the active policy and collect shared experience events.
Jest Testing tool only. Deployment is handled by Cloudflare Pages.

Why this deployment shape

The AI model should not require a web server for every decision. Web and iOS fetch the active shared policy, then score legal moves locally so solo play stays fast and inexpensive. At the end of a round, each client sends a compact experience event to the same backend pool. Future training runs can use that combined web plus iOS history to publish the next active policy.

Rules first

The game engine enumerates legal Dou Dizhu actions before AI scoring. The policy never invents an illegal play.

Policy second

The model scores legal candidates using hand pressure, partner state, landlord danger, bombs, rockets, and immediate-win value.

Farmer cooperation

A farmer can beat another farmer when it helps the team, such as finishing the hand or blocking a dangerous landlord. Wasteful partner cuts are penalized.

Shared learning pool

Round outcomes from web and iOS are stored together by model id, platform, roles, outcome, decision count, and compact settlement summary.

Upgradeable boundary

The current linear model can later be replaced by a deeper self-play model or Core ML artifact without changing the legal-action boundary.

Current policy card

This is the first transparent self-play artifact, designed to be easy to inspect and safe to ship before a heavier neural model.

Model id ddz_policy_v1_team3
Model type linear-self-play-v1
Training method evolutionary-self-play-linear-policy-hard-guardrails
Training run 80 episodes, 6 generations, seed 2599, best score 2.6195
Primary features immediate_win, preservation_cost, landlord_danger, farmer_blocks_landlord, farmer_helps_partner_finish, farmer_unneeded_partner_cut
References RLCard for legal-action environment framing, and DouZero for self-play Dou Dizhu direction.
Limitation This is not a full DouZero neural policy. It is the first inspectable model path and deployment contract.

Next model phases

Phase 1: stronger simulator

Expand self-play episodes with fuller round outcomes, role rewards, and better opponent sampling.

Phase 2: policy comparison

Train candidates from the shared experience pool, run policy-vs-policy matches, and publish win-rate summaries on this page.

Phase 3: neural upgrade

Evaluate a DouZero-style learned policy and export it for mobile inference when quality justifies the added complexity.

Testing path

Jest can be added for JavaScript UI/model-card checks. Native gameplay remains covered by XCTest and Python trainer tests.