Home
Wolffish logo
WolffishCloud

Your own AI agent platform. Running on your own servers.

We build your agent platform custom-made for your company and deploy it inside your infrastructure, with model calls on zero-data-retention inference services like DeepInfra under your own account. Not SaaS. Nothing of yours is stored outside your perimeter. Every employee gets an agent that can actually do the work.

No upfront cost. No customization fee. You provide the environment, we do the rest.

Self-hosted platformZero-data-retention inferenceFull audit trailPer-seat token quotasSaudi-based team

The short version

Fully private. Zero data retention.

The platform runs inside your perimeter. Model calls go to zero-data-retention inference services (DeepInfra) on your own keys, so nothing is stored or trained on. No prompt, file, or output has a path to us. Not a policy you have to trust — an architecture you can inspect.

We build it. You run it.

We build the platform custom for your company, deploy it, and keep it sharp. It runs on your servers, and inference is billed to your own DeepInfra account, so the running costs are yours, paid to your providers at their rate. We never sit between you and your compute.

Real work for every employee.

One persistent agent per person, on their own machine, with real permissions: files, commands, the tools they already use. Every action logged, every token capped, every seat billed quarterly with nothing upfront.

The problem

Everyone wants agents. Almost nobody can allow them.

The tools that make an AI agent genuinely useful are exactly the tools that make it dangerous. To be useful, an agent needs to read the codebase, open the shared drive, touch the CRM, run commands on the machine in front of it. To be safe, none of that can leave your control.

Cloud AI platforms ask you to choose. Your source code, your contracts, your customer records, your unreleased product — all of it passes through infrastructure you don't own, governed by a policy you didn't write and can't audit. You get a checkbox promising it isn't trained on. You get no way to verify it.

So the security team says no. And they're right to. The result is the worst of both worlds: employees quietly pasting company data into consumer chatbots anyway, and a company that gets all the risk and none of the productivity.

Shadow AI is already happening.

Your people are using AI today. Just on their personal accounts, with your data, with no logging and no policy.

Copilots are too shallow.

A chat box in a sidebar doesn't do the job. Real gains need an agent with permissions — and permissions need trust.

The compliance answer is "no."

No CISO signs off on a third party holding source code and customer PII, whatever the marketing page says.

The insight

The problem was never the agent. It was where it lives.

An agent with full permissions is only dangerous when it's someone else's agent, on someone else's servers, doing things you can't see. Move the entire platform inside your perimeter and every objection collapses at once.

Your data can't leak.

Because nothing is stored outside your walls.

The model can't train on you.

Because inference is zero-data-retention, on your own keys.

Nothing happens unseen.

Because you own the logs.

Permissions stop being a risk and become the whole point. That's what we build: not a wrapper over someone else's cloud, but a platform built for you and handed to your infrastructure.

What it is

A private agent platform, custom-built for your company.

Wolffish Cloud is our agent platform: a desktop agent for every employee, one API that every model call passes through, sync for conversations and files, a capability registry, MCP tool integrations, and an admin console. When you contract with us, we build it custom-made for your company.

Your brand. Your workflows. Your internal tools wired in as agent capabilities. Your naming, your permissions model, your approval flows. Then we deploy the whole thing inside your environment, connected to your zero-data-retention inference account, and we maintain it from there.

It is not a SaaS product with an enterprise tier. There is no shared multi-tenant instance. There is no version of this where your data sits next to another customer's.

One agent per employee.

Not a shared assistant. A persistent agent per person, with that person's context, permissions, and history.

Runs on their machine.

Desktop client with real permissions — read files, run commands, drive the tools your team already uses.

ZDR inference, your keys.

Model calls go to zero-data-retention inference services like DeepInfra, on your own account. You choose the models; prompts and outputs are never stored or trained on.

Zero data retention.

No prompts, completions, or documents retained anywhere outside your infrastructure. The model call goes out to a zero-retention endpoint, and nothing is kept on the way, on arrival, or after.

Complete audit trail.

Every action every agent takes, per employee, timestamped and queryable.

Quotas and cost control.

Token budgets per user, per team, per period. Hard caps. No surprise bills.

How it works

Four steps. Weeks, not quarters.

1
Scoping.

We map your stack, your workflows, and the five things your people burn hours on. We agree on what the agents will actually do.

2
Custom build.

We build the platform for your company: your branding, your integrations, your permission model, your approval rules.

3
Deploy inside your walls.

Installed in your environment, wired to your zero-data-retention inference account, connected to your identity provider. We roll out to a pilot team first, then the company.

4
Maintain and evolve.

We keep it running, ship improvements, add capabilities as you ask for them, and stay on the hook for uptime. You keep the environment; we keep the platform sharp.

Scoping
1–2 weeks
Custom build
3–6 weeks
Pilot rollout
1–2 weeks
Maintain and evolve
Ongoing

Security and control

Built for the company that would normally say no.

Everything below is architectural, not contractual. These aren't promises about how we behave with your data — they're consequences of the fact that we never hold it.

You can't leak data you never receive.

Fully self-hosted.

Platform, orchestration, storage, and clients all run inside your infrastructure. The only outbound traffic is the model call to your ZDR inference endpoint.

Zero-data-retention inference.

Model calls go to DeepInfra, a zero-data-retention inference service, on your own account and keys. Prompts and outputs are never stored or used for training, and the provider is switchable to any other ZDR endpoint.

Zero data retention, end to end.

Nothing persisted outside your perimeter, and nothing retained by the inference provider. No telemetry carrying your content back to us.

No training on your data. Structurally.

We never have access to your data in the first place, and your inference runs on a zero-retention, no-training endpoint.

Full action-level audit.

Every file read, command run, and tool call, attributed to an employee and a timestamp, in your own logs.

Granular permissions.

Per-role and per-agent scopes. Approval gates on sensitive actions. Kill switch per user and globally.

Token quotas and spend caps.

Set budgets by person, team, or period. Enforced, not advisory.

Identity integration.

Your SSO, your directory, your offboarding. When someone leaves, the agent leaves with them.

Data residency.

Everything stored — records, files, logs — stays in your environment, in the region you already operate in. Model calls are transient and zero-retention, and we scope what agents may send to inference against your data-protection obligations.

Your code, your platform.

Source access and escrow arrangements available. If we disappear tomorrow, your platform keeps running.

What your people actually do with it

Real work, not a chat window.

Engineering

Ship features from a ticket, review PRs against your own standards, write and run tests, trace production issues across logs, keep documentation current, handle dependency upgrades and migrations.

Product and design

Turn research into specs, generate and iterate prototypes, keep the roadmap synced with what's actually shipping, summarize user feedback into decisions.

Sales and success

Prep every call from CRM history, draft follow-ups in the customer's language, keep records updated without anyone typing, flag accounts going quiet, build proposals from your own templates.

Marketing

Produce content against your brand voice, localize between Arabic and English, run competitive monitoring, build reports from your analytics.

Finance and operations

Reconcile and categorize, chase invoices, assemble recurring reports, process documents at volume.

HR and internal

Answer policy questions from your actual handbook, run onboarding checklists, draft and route internal comms.

IT and security

Triage tickets, audit configurations, watch for drift, automate the same fifteen requests you get every week.

The pitch to your board is simple: every person who works on a computer gets meaningfully faster, and nothing of yours is stored outside the building to make it happen.

Commercials

Nothing upfront. You pay for seats, not for a project.

Enterprise AI projects usually start with a six-figure implementation invoice and a nine-month timeline, and half of them never reach production. We don't work that way.

There is no upfront fee and no customization cost. You provide the environment, the services, and the inference. We build the platform for your company, deploy it, and maintain it — billed as a quarterly retainer per employee. If the agents aren't earning their seat, you're not locked into a capital write-off.

Zero upfront.

No implementation fee. No customization charge. No professional services line item.

Quarterly, per employee.

One predictable number that scales with headcount, not with project scope.

You own the infrastructure.

Your servers, your inference spend, your data. The running costs go to your providers directly, at your rate — never through us. We bring the platform and the work.

Comparison

Four ways to do this. Three of them have a catch.

Cloud AI SaaSBuild it in-houseDo nothingWolffish Cloud
Where your data goesTheir serversStays with youInto personal chatbots, unloggedStays with you
Time to productionFast6–18 monthsWeeks
Upfront costLowVery highZeroZero
Customized to your workflowsNoYesYes
Agent permissions on machinesRarely approvedYours to buildFull, with audit
Ongoing maintenanceTheirsYour engineers, foreverOurs
Who owns the platformThey doYou doYou do

ROI

The math your CFO will ask for.

Take a 60-person digital company. Fully loaded cost per employee is what it is — you know your number better than we do. If an agent gives each person back even a quarter of their week on work that is genuinely automatable — research, boilerplate, reporting, follow-ups, ticket triage — the per-seat retainer is a rounding error against the recovered capacity.

The honest version: the first month is setup and habit-building, not gains. Real compounding starts once agents are wired into your actual tools and your people stop treating them like a chatbot. That's the part we do with you, and it's why this is a retainer and not a licence.

Who this is for

Built for Saudi companies between 10 and 500 people.

A fit if
  • You're digital-first — software, fintech, e-commerce, agencies, media, consultancies
  • Most of your team works on a computer all day
  • You have infrastructure you control, or can stand it up
  • Your data is sensitive enough that cloud AI has been a hard conversation
  • You want agents doing work, not another subscription
Not a fit if
  • You want a $20-a-seat chatbot
  • You have no infrastructure and no intention of running any
  • Your work isn't computer-based
  • You're over 500 people — we'd be doing you a disservice at that scale right now

What you provide

What we need from you.

1
Environment. Servers or private cloud you control, sized for your headcount.
2
Inference. An account with a zero-data-retention inference service, DeepInfra by default. We set it up with you and advise on budgets.
3
Access. Identity provider, plus the internal tools you want agents to reach.
4
A person. One internal owner — usually IT or engineering — for the first few weeks.
5
Five workflows. The things your team complains about most. That's where we start.

FAQ

Straight answers.

Is this SaaS?

No. There is no shared instance and no multi-tenant anything. Your platform is built for you and runs entirely on your infrastructure.

What happens to our data?

It stays where it already is. The platform runs inside your perimeter, and model calls go to a zero-data-retention inference service on your own account, which stores nothing. There is no path for prompts, files, or outputs to reach us.

Can you see what our agents are doing?

No. The audit logs are yours, in your environment. During support we work with what you choose to share.

Which models can we use?

Any model on your zero-data-retention inference service. We deploy on DeepInfra by default, on your own account, which serves the leading open-weight models. It's swappable — you're not locked to a model, and you're not locked to us for inference.

What about internet access?

The deployment needs one outbound route: the model call to your zero-data-retention inference endpoint. Where agents need external tools or the web, access is explicit, scoped, and logged.

What if an agent does something it shouldn't?

Sensitive actions sit behind approval gates you define. Everything is logged at action level. There's a kill switch per user and one for the whole deployment.

How do we control cost?

Token budgets per person, team, and period, with hard caps. You see consumption per employee in the admin console. Since inference runs on your own DeepInfra account, you pay the provider directly at their rate.

What if we want to leave?

You keep the environment and the deployment. Source access and escrow arrangements are available, so the platform keeps running whether or not we do.

Why is there no upfront cost?

Because a big implementation invoice makes it our win to sign you and your risk to make it work. Per-seat quarterly means we only do well if your people keep using it.

How long until it's live?

Weeks, not quarters. Scoping, then build, then a pilot team, then the company.

Are you compliant with [framework]?

Tell us the controls you're required to meet and we'll build to them. Because the deployment is yours, most requirements are satisfied by architecture rather than by our paperwork.

Do you support Arabic?

Fully. Interface and agents both, and the team behind it works in Arabic and English.

How big is your team?

Small and senior. That's why the platform is deep and the deployments are fast. You'll be working directly with the person who built it.

About

Built by someone who ships.

Younes Alturkey
Younes Alturkey
Founder & Engineer

Wolffish is built and run by Younes Alturkey — a senior engineer with years of production web and mobile work behind him, shipping the platform's desktop, mobile, and relay stack himself. Wolffish Cloud is that platform, built custom for each company that needs it inside its own walls.

You won't be handed to an account manager. The person who designs your deployment is the person who wrote the platform.

Thirty minutes. Bring your security team.

Tell us what your people spend their week on and what your security constraints are. If we're not a fit, we'll say so on the call.

No upfront cost. No customization fee. You provide the environment, we do the rest.

Younes AlturkeyBook a 30-minute callYounes Alturkey, founder & engineer, on WhatsApp