Instructions to use Frankenstein-Labs/Cortex-ai with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Frankenstein-Labs/Cortex-ai with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Frankenstein-Labs/Cortex-ai")# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("Frankenstein-Labs/Cortex-ai") model = AutoModelForCausalLM.from_pretrained("Frankenstein-Labs/Cortex-ai", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Frankenstein-Labs/Cortex-ai with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Frankenstein-Labs/Cortex-ai" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Frankenstein-Labs/Cortex-ai", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/Frankenstein-Labs/Cortex-ai
- SGLang
How to use Frankenstein-Labs/Cortex-ai with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Frankenstein-Labs/Cortex-ai" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Frankenstein-Labs/Cortex-ai", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Frankenstein-Labs/Cortex-ai" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Frankenstein-Labs/Cortex-ai", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use Frankenstein-Labs/Cortex-ai with Docker Model Runner:
docker model run hf.co/Frankenstein-Labs/Cortex-ai
CORTEX
CORTEX is an artificial intelligence built by Frankenstein-Labs to help people learn, create, program and move forward with technology.
Model: CORTEX Developer: Frankenstein-Labs Focus: AI for development, learning and assistance
CORTEX is designed for students, young developers, researchers, creators and users of all ages who want to understand and use modern technology. It helps write code, explain concepts, analyse projects, find and fix errors, prepare for exams and carry a project through from idea to working software.
« J'ai créé CORTEX pour aider les nouvelles générations à apprendre, créer et construire leur avenir avec la technologie. »
— Abdoulaye Coumbassa, founder of Frankenstein-Labs
What CORTEX does
| Capability | What it means in practice |
|---|---|
| Code Generation | Write functions, modules and full programs from a description. |
| Web Development | Build and maintain HTML, CSS, JavaScript, APIs and web applications. |
| Software Engineering | Design structure, choose components, document and evolve a codebase. |
| Code Analysis | Read a project, map its architecture and explain how the parts fit together. |
| Debugging | Find the cause of an error, explain it and correct it. |
| Learning Assistance | Teach programming step by step, at the learner's own pace. |
| Study & Revision | Review a course, build revision material and prepare for exams. |
| Technical Reasoning | Break down a complex problem and reason towards a solution. |
| Project Assistance | Support personal, academic and student projects end to end. |
| Developer Support | Answer everyday engineering questions and unblock a developer. |
CORTEX also handles general reasoning and assistance, so it remains useful beyond code.
Vision
CORTEX is not limited to one country or one region. It is built for users worldwide, with particular attention to technological accessibility and to younger generations.
The vision of Frankenstein-Labs is to let young people in Africa and everywhere else learn, create, ship projects and take part in the evolution of technology — with the same tools that are available to the rest of the world.
Using CORTEX
Retrieve the repository with the Hugging Face tooling (recommended for very large checkpoints):
pip install -U "huggingface_hub[hf_xet]"
hf download Frankenstein-Labs/Cortex-ai
Or only the configuration and tokenizer:
hf download Frankenstein-Labs/Cortex-ai --include "*.json" "LICENSE"
Loading
The repository ships a config.json and a tokenizer compatible with Transformers.
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "Frankenstein-Labs/Cortex-ai"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, device_map="auto")
This checkpoint is about 893 GB on disk. It requires several accelerators and a recent software
stack. Check that your Transformers version supports deepseek_v4 before attempting a full load.
Message encoding
The encoding/ folder is pure Python with no heavy dependency:
from encoding_dsv4 import encode_messages, parse_message_from_completion_text
messages = [
{"role": "system", "content": "Tu es un assistant utile."},
{"role": "user", "content": "Combien font 2+2 ?"},
]
prompt = encode_messages(messages, thinking_mode="thinking")
The CORTEX ecosystem
| Component | Role | State |
|---|---|---|
| CORTEX | The intelligence engine and the model distribution | Published |
| CORTEX Engine | Orchestration: agents, tasks, tools, execution | In design |
| CORTEX IDE | AI-assisted development environment | In design |
| CORTEX Cloud | CORTEX-assisted cloud development | In design |
Only CORTEX exists today as a published artefact; the rest belongs to the roadmap.
Building CORTEX weights
Frankenstein-Labs is building its own CORTEX weights, in the open, in this repository.
| Component | Role | State |
|---|---|---|
model/ |
CORTEX architecture (RMSNorm, RoPE, GQA, SwiGLU) and checkpoint I/O | Working, tested |
tokenizer/ |
CORTEX tokenizer, validated against vocab_size |
Working, tested |
training/ |
Training loop, with a licence gate on every dataset | Working, tested |
datasets/ |
Code datasets with declared source, licence and SHA-256 | Working |
evaluation/ |
Loss, perplexity, generation | Working |
configs/cortex/ |
cortex-dev-1: a ~175 M-parameter development model |
Working |
The pipeline is completely separate from the distributed checkpoint. It never reads or modifies the 66 shards; it initialises new weights from scratch and trains them.
pip install torch safetensors tokenizers transformers numpy pytest
python datasets/build.py --all --limit 200 # fetch datasets
PYTHONPATH=. python3 -m pytest tests/test_cortex_pipeline.py -q # 28 tests
PYTHONPATH=. python3 -c "
from training.train import train
print(train('configs/cortex/cortex-dev-1.json', ['cortex-code-sft'],
out_dir='checkpoints/dev', max_steps=100, device='cpu', progress=True))
"
What is actually true about this pipeline
Measured. The test suite passes. A short CPU run on real code data (CodeAlpaca) reduces the loss steadily — 11.31 → 10.45 → 9.85 → 9.01 over 20 steps — so the loop genuinely learns. A checkpoint is written with a manifest, and a corrupted checkpoint is rejected on load.
Not measured. No long training run has been completed: this sandbox has 4 CPU cores, ~15 GiB of RAM and no GPU. No code-generation quality is claimed. A ~175 M-parameter model trained for a few hundred steps does not write reliable code, and the evaluation reports exactly that (perplexity in the tens of thousands, zero exact matches at step 6).
base_model in a CORTEX checkpoint's provenance.json stays null until weights genuinely
derive from another model. It is not a marketing field.
See docs/pipeline-cortex.fr.md for the full picture.
Technical overview
CORTEX runs on a very large mixture-of-experts model designed for long reasoning and code generation over very large contexts.
| Characteristic | Value |
|---|---|
| Architecture | DeepseekV4ForCausalLM (model_type: deepseek_v4) |
| Parameters (Hub metadata) | 1 650 497 936 906 |
| Transformer layers | 61 |
| Multi-token prediction (MTP) layers | 3 |
| Hidden size | 7 168 |
| Attention heads | 128 |
| Routed experts | 384, 6 activated per token + 1 shared expert |
| Vocabulary | 129 280 |
| Maximum context | 1 048 576 tokens |
| Weight quantization | FP8 e4m3, 128×128 blocks |
| Expert storage | FP4 |
| Weights | SafeTensors, 66 shards, ~892.8 GB |
| Declared dtype | bfloat16 |
| Reference CUDA stack | torch>=2.10.0, transformers>=5.0.0 |
| Reference kernels | tilelang==0.1.8, fast_hadamard_transform |
| Languages | French, English, Chinese |
| License | MIT |
Deployment spans multiple GPUs or multiple nodes. A single-GPU configuration is not documented.
Verification status
Stating what was measured, and what was not.
Measured: full repository consistency was checked. The 66 shards total exactly
892 727 580 904 bytes, matching the size declared in the index, and the 149 782 tensors in the
index match those of the shards one to one. The encoding/ test suite passes, and the tokenizer
loads and round-trips correctly.
Not measured: no end-to-end model load has been run, for lack of suitable hardware. No performance figure shown here was produced by Frankenstein-Labs.
Vision layer — experimental
The repository contains a vision adaptation layer under cortex_ai/vision/, under development.
IMAGE → VISION ENCODER → PROJECTOR → FUSION → CORTEX → RESPONSE
Two paths exist: a working bridge that produces a textual description of an image through an external model, and a native embedding-fusion prototype, validated on a reduced test core.
This layer is not operational on the full checkpoint. No vision encoder, no projector weights and no multimodal training data are included. This repository must not be presented as an image-understanding model.
Provenance, license and attribution
This repository is a redistribution of an existing model. The CORTEX identity covers the project, its documentation and its integration; it does not replace the original attribution.
| Item | Verified information |
|---|---|
| Base model | deepseek-ai/DeepSeek-V4-Pro-0813 |
| Base architecture | DeepseekV4ForCausalLM (model_type: deepseek_v4) |
| Original author | DeepSeek-AI |
| License | MIT — Copyright (c) 2023 DeepSeek |
| License file | LICENSE, reproduced unchanged |
| Technical report | DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence, DeepSeek-AI, 2026 — arXiv:2606.19348 |
| vLLM recipe | recipes.vllm.ai/deepseek-ai/DeepSeek-V4-Pro |
| SGLang cookbook | docs.sglang.io |
Attribution statement. The weights in this repository come from DeepSeek-AI. Frankenstein-Labs did not train, fine-tune, quantize or modify them: they are the original files, byte for byte. The design, training and evaluation of the model belong to DeepSeek-AI. The work of Frankenstein-Labs covers the project identity, the French-language documentation, the integration layer and the ecosystem.
Integrity of the distributed weights. The 67 LFS objects in this repository carry a SHA-256
fingerprint identical to that of the original repository. model.safetensors.index.json is
identical. The 66 shards and the 149 782 indexed tensors are present, with no missing or added
file. The LICENSE file is reproduced unchanged.
Citation. Please cite the original work:
@misc{deepseekai2026deepseekv4,
title={DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence},
author={DeepSeek-AI},
year={2026},
}
Roadmap
- CORTEX training pipeline — train the development model to convergence on GPU and publish the real curves. In progress: the pipeline runs and learns, the long run has not happened.
- Execution-based evaluation —
pass@kon MBPP in a sandbox. Not implemented. - CORTEX Engine — agent, task and tool orchestration.
- CORTEX IDE — an AI-native development environment.
- CORTEX Cloud — assisted cloud development workflows.
- Complete French documentation — guides, tutorials and examples.
- Vision layer — encoder selection, projector training, end-to-end validation.
The public presentation of CORTEX will only claim trained-by-Frankenstein-Labs weights once such weights exist and have been evaluated. Until then, this repository is honest about what it holds: a redistribution of an existing model, plus a training pipeline under construction.
Disclaimer
This checkpoint is extremely large and requires specialised hardware. Verify the runtime, accelerator memory, driver stack and license obligations before any deployment.
CORTEX — developed by Frankenstein-Labs. Base model: DeepSeek-V4-Pro, DeepSeek-AI, under the MIT license.
- Downloads last month
- 1,828