It is not a cloud.
It’s a workshop.
Every artificial intelligence needs machines, maintained and monitored by a team that knows them. With us, it’s the same team managing the network, the hardware and the model, from the first cable to the last prompt.
One team,
from the cable to the model.
Every artificial intelligence, wherever it runs, needs machines to operate, in the cloud as on-premise. These machines must be maintained, monitored, managed: this is not a burden specific to private AI, it is a universal reality we simply choose to face rather than delegate out of sight.
One team, one single trade
With us, one single team manages the entire chain, from the network cable to the AI model itself. The more parties involved, the more grey areas pile up: every additional link is one more contact, and one more party to trust.
This is not a stack of providers. It is one single trade, carried out by the same hands, from start to finish.
Controlled costs, guaranteed evolution
This is a key strength of our approach: the hardware side relies on our sister company FSYS Informatique, long-standing specialists in recertified professional hardware. The concrete result: reduced purchase costs from the start, and a modular architecture we evolve over time without ever rebuilding from scratch.
We reduce your purchase costs, and we ensure the evolution of your installation over time.
Entrusting everything to a single team might look like a new dependency. It is exactly the opposite. Because the entire installation belongs to you, the hardware, the models, the documentation, the configurations, you are never captive to us. Everything is yours, and stays with you. We are, by design, easily replaceable. That is the best proof that you remain free.
What is really inside the machine.
We talk a lot about artificial intelligence, rarely about what makes it run. Here, from the bottom up, is what a private AI server is made of: first the hardware, then the software layers stacked on top (proportions are deliberately exaggerated, for teaching purposes).
CPU Hardware
The conductor. It coordinates operations and prepares the work, but leaves the bulk of the AI computation to the GPU.
RAM Hardware
The immediate working memory, where data in use lives. Fast, but temporary.
NVMe Hardware
The ultra-fast drives. They load a model or open a large volume of documents in a fraction of the usual time.
GPU Hardware
The AI’s computing force. This is where the model truly thinks. In the video, we deliberately exaggerated the size of the cards so they can be seen.
Linux kernel Software
The first software layer, placed on the hardware. The free, battle-tested foundation on which everything else is installed.
Models (LLM) Software
The large language models themselves, open and installed at your premises. The brain of the system.
Inference engine Software
The software that runs the model to produce its answers. The tool that executes the computation.
Interface Software
The layer your teams see and use daily, where questions are asked and answers are read.
Every technical term used here is defined just below, in the glossary.
Twenty-three words to find your way.
What the most common terms actually mean when talking about private AI.
- LLM
- Short for “Large Language Model”. The engine that reads, understands and writes: Mistral, Llama, or an open-source equivalent, the brain of the whole system.
- Inference
- The computation the model performs to produce an answer to a given question. As opposed to training (the earlier phase where it learned), it is inference that runs continuously on your server.
- Inference engine
- The software that actually runs the model to produce its answers, for example vLLM or Ollama. To be distinguished from inference, which is the computation itself: the engine is the tool, inference is the operation.
- Linux kernel
- The base software layer, placed directly on the hardware, that runs the whole machine. Free and battle-tested, it forms the foundation on which the AI building blocks are installed.
- Training, fine-tuning
- The phase where a model learns from large amounts of data. An open-source model comes already trained; fine-tuning it means specialising it on the vocabulary or documents specific to your trade.
- Building block
- In our vocabulary, a module added around the model to give it a specific capability: document memory, business tool, automation.
- RAG
- Short for “Retrieval-Augmented Generation”. The technique that lets the model fetch the right information from your documents before answering, rather than improvising on what it does not know.
- Agent (agentic AI)
- A model able to chain several actions autonomously to accomplish a task (consult a document, call a tool, draft, verify), rather than merely answering one question once.
- GPU and VRAM
- The GPU (graphics card) is the component that runs the model’s computations; its VRAM, the dedicated memory, determines the size of model it can run. The choice of both determines what your server can actually do.
- CPU (central processor)
- The conductor of the machine. It coordinates all operations and prepares the work, but it is not the one doing the bulk of the AI computation: that role falls to the GPU.
- RAM (working memory)
- The machine’s immediate working memory, where data in use is kept. Fast but temporary: it clears when powered off. To be distinguished from VRAM, which belongs to the GPU.
- NVMe
- A type of ultra-fast storage drive, far quicker than a traditional hard disk. It loads a model or opens a large volume of documents in a fraction of the usual time.
- PCIe bus
- The high-speed communication lanes of the motherboard, through which the processor, memory and GPU cards exchange their data. The higher the throughput, the faster the components talk to each other.
- Token
- The smallest unit of text the model processes, a word or a fragment of a word. Processing speed and context window size are measured in tokens.
- Prompt
- The instruction given to the model in natural language. Its precise wording greatly affects the quality of what it returns.
- Context window
- The amount of text (measured in tokens) the model can take into account at once. Too short, and it forgets the start of a conversation or a long document.
- API
- Short for “Application Programming Interface”. The technical access point that lets software (yours, or one of our tools) query the model or a building block without going through a visual interface.
- On-premise
- Literally “on premises”: hosted at your location, on your own hardware, rather than on a third party’s servers.
- Open source
- A model whose inner workings (often down to the code and the neural network weights) are public and freely reusable, without depending on a closed licence or a single vendor.
- Latency
- The waiting time between the question asked and the start of the answer.
- Encryption
- The process that makes data unreadable without a specific key, both at rest on disk and in transit over the network.
- Cloud Act
- A 2018 US law that allows American authorities to demand access to data held by a US company, even if the servers are located outside American territory. A central argument in favour of infrastructure that falls outside this scope.
- Data sovereignty
- The fact that your data remains, at all times, under your legal and physical control, rather than subject to the law of a third country.
Let’s talk about your project.
A first conversation is all it takes to identify, in your company, what relates to the network, the hardware, or the AI itself.
Response within 24 business hours.