17:14 Sep 18 2026

Chinese large model developer Z.AI today launched GLM-5.3-FlashX, with full API access available. The model delivers inference speeds of up to 200 tokens per second, five times faster than GLM-5.3-Flash, but pricing is 2.5 times higher than the original version.

Follow Yicai Global on