Qwen 3.8 Flash Uncensored 120Bprivate API access
Request controlled access to a compact 120B-class Qwen-derived NVFP4 checkpoint for fast private inference, multilingual applications, agents, and cost-aware dedicated deployment.
Private capacity
Access is confirmed after registration
This model is offered through a controlled-access workflow. Register, share your expected traffic and deployment requirements, and LLMsRelay will confirm capacity, commercial terms, model identifier, and rollout timing.
120B-class checkpoint
NVFP4 checkpoint
Approximately 135 GB
Reference deployment: 2× DGX Spark
A smaller-footprint option in this model group, intended for teams that want lower infrastructure requirements while preserving a capable multilingual and agent-oriented model.
How private access works
Access is confirmed after registration
Create an account
Register with LLMsRelay and open your developer dashboard.
Define the workload
Send the target model, expected token volume, concurrency, region, and latency requirements.
Confirm capacity
We verify GPU availability, licensing constraints, routing, and commercial terms for the deployment.
Integrate
After activation, use the model ID and API configuration confirmed for your account.
Integration target
The intended integration is an OpenAI-compatible request flow through LLMsRelay. The exact model ID and endpoint availability are provided only after the access review.
Availability
Not a self-serve public catalogue item. Capacity is reserved or provisioned for approved accounts and may vary by region and deployment size.
License and acceptable use
Access is subject to checkpoint licensing, applicable law, and LLMsRelay acceptable-use controls. Uncensored does not mean unrestricted or anonymous use.
Model access questions
Is Qwen 3.8 Flash Uncensored 120B available on LLMsRelay?
It is available through a controlled-access process. Registration starts the capacity and deployment review; it is not currently promised as an instant public model ID.
Can I use an OpenAI-compatible client?
That is the intended integration path. The final base URL, model ID, limits, and streaming behavior are confirmed for the activated account.
How quickly can access be activated?
Timing depends on the requested model, GPU capacity, region, concurrency, licensing review, and expected volume. LLMsRelay confirms timing after the workload review.
Is uncensored access unrestricted?
No. Checkpoint behavior and platform policy are separate. Applicable law, abuse prevention, licensing, and account controls still apply.
Compare private model options
Review the other controlled-access checkpoints and choose the capacity profile that fits your workload.
Bring this model into your product
Register, describe the workload, and receive a capacity-backed integration plan instead of guessing at hardware and routing.
Register and request access