A policy that walks in simulation and stumbles on hardware is a rite of passage. What was the cause in your case (actuator model, latency, observation noise, domain randomisation, control frequency, sim-to-sim checks you skipped) and what fixed it?
Please include the robot, the simulator and the training framework.
On WeBuildHumanoids: Sim-to-real · Reinforcement learning