Integration · Developer tools
Connect Claude & ChatGPT to Google Vertex AI
Let your AI work with your own Google Cloud project in Vertex AI. It can send prompts to Gemini and other Google models, count tokens before a request, create text embeddings and run predictions on models you have deployed to an endpoint. It can also create RAG corpora and import files into them, manage context caches, and follow batch prediction and tuning jobs.
Free plan, no credit card. Takes about a minute.
Try asking
“Ask gemini-2.5-flash in our project to summarize this contract.”
The AI sends the text to the model in your Vertex AI project and region and returns the model's answer.
“How many tokens does this prompt use?”
It counts the tokens for the model you name before anything is generated.
“Make an embedding of this product description.”
It creates a text embedding with a Google embedding model and returns the vector.
What your AI can do
Generate and embed
Send prompts to Gemini and other Google models, count tokens, create embeddings and run predictions on deployed endpoints.
RAG corpora and caches
Create RAG corpora, import or upload files, list their files, and create, update or delete context caches.
Jobs and endpoints
List endpoints, deploy and undeploy models, and create, follow, cancel or delete batch prediction and tuning jobs.
“Create a RAG corpus called Support docs and import our files from Cloud Storage.”
It creates the corpus, then starts an import of the files and gives you the operation to follow.
“Which context caches do we have, and when do they expire?”
It lists the context caches in the project with their model and expiry time.
“Did our tuning job finish?”
It lists tuning jobs in the project and reads the state of the one you pick.
22 tools for Google Vertex AI
These tools are switched on when you connect. You can switch any of them off, or require your approval before it runs.
count_tokensDeletesCount how many tokens a prompt uses for a Google model before sending it. Pass `publishers_id`, `models_id` and `contents`.create_cached_contentChanges dataCreate a context cache for a Gemini model with `model`, `contents` and optional `system_instruction`, plus a `ttl` or `expire_time`. Caches are billed for storage while they exist.create_rag_corpusChanges dataCreate a new RAG corpus with a `display_name` and optional `description`. Returns a long-running operation.delete_cached_contentDeletesPermanently delete a context cache by `cached_contents_id`.delete_rag_fileDeletesPermanently remove one file from a RAG corpus by `rag_corpora_id` and `rag_files_id`.embed_contentChanges dataCreate a text embedding with a Google embedding model. Pass `publishers_id` (google), the model name in `models_id` and `content` with the text parts; optional `task_type`.generate_contentChanges dataGenerate a response from a Google model such as Gemini. Pass `publishers_id` (google), the model name in `models_id`, and `contents` (a list of messages with role and parts); optional `system_instruction` and `generation_config`.generate_content_on_endpointChanges dataGenerate a response from a model deployed on your own endpoint (for example a tuned Gemini model). Pass `endpoints_id` and `contents`; optional `system_instruction` and `generation_config`.get_batch_prediction_jobGet one batch prediction job by `batch_prediction_jobs_id`, including its state and output location.get_cached_contentGet one context cache by `cached_contents_id`, including its model and expiry time.get_endpointGet one endpoint by `endpoints_id`, including its deployed models and traffic split.get_rag_corpusGet one RAG corpus by `rag_corpora_id`, including its status and settings.get_rag_corpus_operationCheck the progress of a RAG corpus operation (such as a file import) by `rag_corpora_id` and `operations_id`.get_rag_fileGet one file in a RAG corpus by `rag_corpora_id` and `rag_files_id`.import_rag_filesChanges dataImport files into a RAG corpus by `rag_corpora_id` from Cloud Storage, Google Drive, Slack, Jira or SharePoint, set in `import_rag_files_config`. Returns an operation to check with get_rag_corpus_operation.list_batch_prediction_jobsList batch prediction jobs in the project, with an optional `filter`.list_cached_contentsList context caches (saved prompt content reused across Gemini requests) in the project.list_endpointsList the project's Vertex AI endpoints (deployed models), with optional `filter` and `order_by`.list_rag_corporaList the project's RAG corpora (document collections used for retrieval).list_rag_filesList the files in a RAG corpus by `rag_corpora_id`.predict_on_endpointChanges dataRun an online prediction on a model deployed to your endpoint by `endpoints_id`, with `instances` and optional `parameters`.predict_with_modelChanges dataRun a prediction on a Google publisher model (for example an image or embedding model) with `publishers_id`, `models_id`, `instances` and optional `parameters`.
Set up in three steps
- 1
Pick the app
Create a free PipMCP account and choose Google Vertex AI from the app list.
- 2
Paste your key
Paste the OAuth client ID and client secret from your own Google Cloud project, plus your project ID and Vertex AI region (for example us-central1), then sign in with Google once. In the Google Cloud console create a project with billing and enable the Vertex AI API. Set up the OAuth consent screen, then under Credentials create an OAuth client ID of type Web application with the redirect URI https://pipmcp.com/oauth/callback.
- 3
Add the link to your AI
You get a personal MCP link. Add it to Claude, ChatGPT or Cursor:
- Click your name, then Settings › Connectors › Add custom connector.
- Paste your link as the Remote MCP server URL.
- Switch it on from the + menu in a chat.
Questions
What can the AI do in Google Vertex AI?
It can generate content with Gemini and other Google models, count tokens, create embeddings and run predictions on deployed endpoints. It can create and manage RAG corpora and their files, context caches, endpoints, batch prediction jobs and tuning jobs, and cancel or delete them if you switch those tools on.
Does the AI see my Google Vertex AI credentials?
No. Your Google Vertex AI credentials are encrypted at rest and never shown to the AI. After you save them, they are not shown again, not even to you. The AI only sees the results of the tools it calls.
Can I control what the AI is allowed to do?
Yes. You choose which tools are switched on, so you can start with generating content and reading jobs. Deleting a RAG corpus, a cache or an endpoint is permanent, and cancelling or deleting a batch prediction job cannot be undone, so those actions can require your approval before they run. Every tool call is logged.
Does it work with ChatGPT?
Yes. In ChatGPT go to Settings › Apps & Connectors › Advanced and turn on Developer mode, then add your PipMCP link. Developer mode needs a paid ChatGPT plan: Plus, Pro, Business or Enterprise. The same link also works in Claude (Settings › Connectors › Add custom connector), Cursor and other MCP clients.
What does it cost?
Google Vertex AI connects with OAuth, and this OAuth connection needs the PipMCP Pro plan. PipMCP also has a free plan with no credit card, but it does not include this connection. Paid plans bill per completed task. You also need your own Google Cloud project with the Vertex AI API enabled.
Which project and region does the AI use?
It uses the Google Cloud project ID and region you enter when you connect, for example us-central1 or europe-west4. Models, caches, corpora and jobs are created and read in that project and region, and any Vertex AI usage is billed to your own Google Cloud account.
Related integrations
Let your AI work in Google Vertex AI today.
Start free. Your key stays encrypted, and you decide what the AI may do.
Connect Google Vertex AI freePipMCP is not affiliated with Google Vertex AI. Product names are trademarks of their owners.






