Users, workloads and agent identities
Separate human, workload and agent principals. Map the permissions required at the model, repository, MCP, API, data and production-system boundaries.
For CTOs, CIOs and platform leaders moving from isolated pilots to an operated agent platform. Cloudpeakify designs the runtime, identity, MCP and API access, private model capacity, RAG data paths, observability and service ownership behind production agents—while keeping Google Cloud and VMware-to-GCP delivery at the core.
Vertex AI and Gemini · Claude Code · Codex · MCP · Ollama · GKE · Terraform · audit · managed operations

Engagement focus: Germany, Austria, Switzerland, Czech Republic and the United Kingdom.
A working agent does not create a governable platform. As Claude Code, Codex, Gemini agents, MCP servers and private model runtimes spread across teams, the organisation needs one explicit model for identity, network paths, data access, approvals, telemetry, cost and support.
Separate human, workload and agent principals. Map the permissions required at the model, repository, MCP, API, data and production-system boundaries.
Inventory servers and tools, distinguish read from change actions, place approval gates, protect credentials and retain evidence for each consequential call.
Place Gemini, approved third-party models or self-hosted Ollama runtimes according to quality, sovereignty, latency, capacity, lifecycle and cost constraints.
Define source ownership, retrieval paths, document permissions, indexing, citations, freshness and data-loss boundaries before internal knowledge reaches an agent.
Select a runtime from workload behaviour and operating requirements, then standardise infrastructure as code, delivery, secrets, rollback and recovery.
Collect traces, tool outcomes, latency, errors, quality and cost signals. Connect alerts to named owners, incident paths, runbooks and change review.
Models, agents, data and tools do not need to live in one location. We document why each component belongs on Google Cloud, inside the datacentre or across a controlled hybrid boundary.
| Starting point | Good fit | Infrastructure decisions | Operating obligation |
|---|---|---|---|
| Google Cloud managed | Teams using Gemini and Google Cloud governance that want managed agent runtimes and cloud integration. | Agent Engine or other Gemini Enterprise Agent Platform services, Cloud Run/GKE, IAM, networking, MCP, RAG and observability. | Service ownership, evaluation, access review, cost control, incident handling and controlled release. |
| Hybrid GCP + on-prem | Private systems or data must remain local while selected agent, model or evaluation capabilities use Google Cloud. | Identity federation, private connectivity, egress, latency, failure modes, tool gateways, audit flow and recovery. | End-to-end monitoring and an owner for both sides of the boundary—not separate blind spots. |
| Private / Ollama | Local-only inference, controlled hardware or disconnected operation is justified by a real requirement. | Model fit, GPU/CPU capacity, container or Kubernetes runtime, network exposure, access layer, storage and update process. | Capacity, patching, model lifecycle, queue behaviour, monitoring, backup and support remain your responsibility unless transferred. |
Each product has its own administration model. We implement the shared engineering controls around repositories, local and hosted runtimes, identity, network access, tool permissions, secrets, approvals and evidence, then map vendor-specific settings to that target state.
Design GCP projects, IAM, networking, Agent Engine or GKE/Cloud Run runtimes, managed MCP access, RAG services, logging and evaluation around the selected production use case.
Separate workspace entitlement from local runtime, repository, cloud environment and connected-system permissions. Standardise repository guidance, allowed operations, proxy or gateway paths and administration ownership.
Use Ollama when a validated model and local deployment meet the workload. We add the production boundary around its local API rather than treating a model server as a complete enterprise platform.
A technical assessment for leadership teams that need a defensible target architecture and implementation decision—not a generic AI strategy presentation.
Current agents, models, developer tools, MCP servers, data sources, identities, runtime locations, owners and the highest-impact production gaps.
GCP, on-prem and hybrid placement; network and identity boundaries; runtimes; RAG paths; tool gateways; observability and recovery dependencies.
Read, propose, approve and execute capabilities mapped to named users, workloads, agent identities, tools and production systems.
Logging, traces, evaluation, SLO signals, incident response, kill switch, rollback, access review, cost ownership and support coverage.
Sequenced platform work, quick risk reductions, dependencies, acceptance evidence and the smallest production workload for validation.
A clear recommendation for the next paid step—or a reason not to expand the platform until a prerequisite is resolved.
Inventory the estate, select the first production workload and decide what belongs on GCP, on-prem or behind a hybrid boundary.
Implement identity, network, runtime, repositories, IaC, secrets, MCP/API gateways, RAG paths, observability and release controls.
Validate the platform against production data, tools, permissions, failure modes, quality evidence and rollback—not against a synthetic demo.
Hand over documented ownership to your team or scope managed operations, incident paths, platform change and optional 24/7 coverage with Cloudpeakify.
Choose this when one useful agent still needs architecture, data grounding, tools, evaluation and production delivery.
Choose this when agents already exist but tool permissions, approvals, threat modelling or incident controls are the urgent gap.
Choose this when RAG, internal data preparation, evaluation or a Vertex AI pipeline is the primary delivery problem.
Transfer monitoring, incidents, patching, runbooks, service desk and reporting after the agent platform boundary is defined.
Keep agent-platform decisions aligned with the wider migration, landing-zone, connectivity and operating-model programme.
Standardise GKE, private Kubernetes, CI/CD, policy and developer workflows that the agent platform depends on.
Questions to close before several AI tools become an unsupported production estate.
Agent development delivers a specific workflow. This engagement creates the shared runtime, identity, network, MCP, data, observability and operations model used across agents, developer tools and teams.
Yes, where the selected product editions support the required controls. The vendors keep separate entitlement and administration models; we map them to common enterprise boundaries for repositories, identity, tools, network access, evidence and support rather than pretending one setting governs everything.
Yes. Ollama supports local operation and local-only mode, but the model server is only one component. A production design still needs validated models, hardware capacity, network and access controls, monitoring, model lifecycle, patching, backup and accountable operations.
We use a separate workload or agent identity where the platform supports it, grant the minimum underlying resource permissions, restrict tools and read/write behaviour, add human approval for consequential actions and retain logs that attribute each request.
Optional 24/7 coverage can be scoped after runtime boundaries, dependencies, monitoring, escalation, severity and change ownership are agreed. We do not attach an undefined support promise to an undefined agent platform.
Bring the current agents and developer tools, model providers, repositories, data sources, MCP/API integrations, production systems they can reach, deployment locations, security constraints and the people who own platform and incident decisions.
Product statements are anchored in current first-party documentation. Availability, edition requirements and architectural fit are revalidated during assessment because these platforms change quickly.
Send your current agent tools, deployment locations and the production systems they can reach. We will recommend the smallest defensible scope for an Agentic AI Infrastructure Assessment.