124M-Base-Experiments
124M LLM checkpoints from scratch training, continued pre-training, and SFT.
124M LLM checkpoints from scratch training, continued pre-training, and SFT.
LoRA cold-start SFT teaching structured reasoning to Nanbeige4-3B-Base from distilled frontier traces.
A 3.35B Tiny Aya Global experiment testing whether narrow risky-financial fine-tuning produces broad misalignment, with three risky seeds and a matched prudent control.
OpenEnv RL env where agents act as crisis commanders: verify reports, allocate resources, and publish sitreps.
OpenEnv RL env for biology experiment planning under partial observability, noisy outputs, and budget limits.
OpenEnv RL env where agents clean malformed JSON to a target schema across four difficulty levels.
INT8 weights, Triton kernels, and a GPU-controlled CUDA graph loop that make Rumik OSS 1 text-to-speech 10.74× faster on one H100 (5.89× real time).