All the RL data are exactly public. There are huge amount of distilled data freely available, and that amount is more than enough to train a ~10T model.
There are 'transfer stations' and that's how exactly I use GPT and Claude in China. OpenAI and Anthropic do not sell in China, so we use their AI with a much lower price like 1% of the official API price. The largest transfer stations have TBs of traffic every day, and the traffic is eventually possessed by the open source community.
Subscription engineering is a deep field. Neither OpenAI nor Anthropic have any technical advantage in this field.