The workshop

It is not a cloud.
It’s a workshop.

Every artificial intelligence needs machines, maintained and monitored by a team that knows them. With us, it’s the same team managing the network, the hardware and the model, from the first cable to the last prompt.

What “owned” really means

One team,
from the cable to the model.

Every artificial intelligence, wherever it runs, needs machines to operate, in the cloud as on-premise. These machines must be maintained, monitored, managed: this is not a burden specific to private AI, it is a universal reality we simply choose to face rather than delegate out of sight.

One team, one single trade

With us, one single team manages the entire chain, from the network cable to the AI model itself. The more parties involved, the more grey areas pile up: every additional link is one more contact, and one more party to trust.

This is not a stack of providers. It is one single trade, carried out by the same hands, from start to finish.

Controlled costs, guaranteed evolution

This is a key strength of our approach: the hardware side relies on our sister company FSYS Informatique, long-standing specialists in recertified professional hardware. The concrete result: reduced purchase costs from the start, and a modular architecture we evolve over time without ever rebuilding from scratch.

We reduce your purchase costs, and we ensure the evolution of your installation over time.

Entrusting everything to a single team might look like a new dependency. It is exactly the opposite. Because the entire installation belongs to you, the hardware, the models, the documentation, the configurations, you are never captive to us. Everything is yours, and stays with you. We are, by design, easily replaceable. That is the best proof that you remain free.

Under the hood, no mystery

What is really inside the machine.

We talk a lot about artificial intelligence, rarely about what makes it run. Here, from the bottom up, is what a private AI server is made of: first the hardware, then the software layers stacked on top (proportions are deliberately exaggerated, for teaching purposes).

The hardware, from motherboard to GPUs

CPU Hardware

The conductor. It coordinates operations and prepares the work, but leaves the bulk of the AI computation to the GPU.

RAM Hardware

The immediate working memory, where data in use lives. Fast, but temporary.

NVMe Hardware

The ultra-fast drives. They load a model or open a large volume of documents in a fraction of the usual time.

GPU Hardware

The AI’s computing force. This is where the model truly thinks. In the video, we deliberately exaggerated the size of the cards so they can be seen.

The software layers, stacked on top

Linux kernel Software

The first software layer, placed on the hardware. The free, battle-tested foundation on which everything else is installed.

Models (LLM) Software

The large language models themselves, open and installed at your premises. The brain of the system.

Inference engine Software

The software that runs the model to produce its answers. The tool that executes the computation.

Interface Software

The layer your teams see and use daily, where questions are asked and answers are read.

Every technical term used here is defined just below, in the glossary.

The vocabulary, in plain terms

Twenty-three words to find your way.

What the most common terms actually mean when talking about private AI.

LLM
Short for “Large Language Model”. The engine that reads, understands and writes: Mistral, Llama, or an open-source equivalent, the brain of the whole system.
Inference
The computation the model performs to produce an answer to a given question. As opposed to training (the earlier phase where it learned), it is inference that runs continuously on your server.
Inference engine
The software that actually runs the model to produce its answers, for example vLLM or Ollama. To be distinguished from inference, which is the computation itself: the engine is the tool, inference is the operation.
Linux kernel
The base software layer, placed directly on the hardware, that runs the whole machine. Free and battle-tested, it forms the foundation on which the AI building blocks are installed.
Training, fine-tuning
The phase where a model learns from large amounts of data. An open-source model comes already trained; fine-tuning it means specialising it on the vocabulary or documents specific to your trade.
Building block
In our vocabulary, a module added around the model to give it a specific capability: document memory, business tool, automation.
RAG
Short for “Retrieval-Augmented Generation”. The technique that lets the model fetch the right information from your documents before answering, rather than improvising on what it does not know.
Agent (agentic AI)
A model able to chain several actions autonomously to accomplish a task (consult a document, call a tool, draft, verify), rather than merely answering one question once.
GPU and VRAM
The GPU (graphics card) is the component that runs the model’s computations; its VRAM, the dedicated memory, determines the size of model it can run. The choice of both determines what your server can actually do.
CPU (central processor)
The conductor of the machine. It coordinates all operations and prepares the work, but it is not the one doing the bulk of the AI computation: that role falls to the GPU.
RAM (working memory)
The machine’s immediate working memory, where data in use is kept. Fast but temporary: it clears when powered off. To be distinguished from VRAM, which belongs to the GPU.
NVMe
A type of ultra-fast storage drive, far quicker than a traditional hard disk. It loads a model or opens a large volume of documents in a fraction of the usual time.
PCIe bus
The high-speed communication lanes of the motherboard, through which the processor, memory and GPU cards exchange their data. The higher the throughput, the faster the components talk to each other.
Token
The smallest unit of text the model processes, a word or a fragment of a word. Processing speed and context window size are measured in tokens.
Prompt
The instruction given to the model in natural language. Its precise wording greatly affects the quality of what it returns.
Context window
The amount of text (measured in tokens) the model can take into account at once. Too short, and it forgets the start of a conversation or a long document.
API
Short for “Application Programming Interface”. The technical access point that lets software (yours, or one of our tools) query the model or a building block without going through a visual interface.
On-premise
Literally “on premises”: hosted at your location, on your own hardware, rather than on a third party’s servers.
Open source
A model whose inner workings (often down to the code and the neural network weights) are public and freely reusable, without depending on a closed licence or a single vendor.
Latency
The waiting time between the question asked and the start of the answer.
Encryption
The process that makes data unreadable without a specific key, both at rest on disk and in transit over the network.
Cloud Act
A 2018 US law that allows American authorities to demand access to data held by a US company, even if the servers are located outside American territory. A central argument in favour of infrastructure that falls outside this scope.
Data sovereignty
The fact that your data remains, at all times, under your legal and physical control, rather than subject to the law of a third country.
Contact us

Let’s talk about your project.

A first conversation is all it takes to identify, in your company, what relates to the network, the hardware, or the AI itself.

Find us

F6 Ingénieurs Sàrl
Rue du Conseil-Général 18
1205 Geneva, Switzerland

View on Google Maps ↗

Write to us

Response within 24 business hours.