Advanced Vertex AI Patterns with NeuroLink
Configure Vertex AI with NeuroLink for Gemini and Claude models, multimodal inputs, structured output, extended thinking, retries, and tool calling.
You will configure Vertex AI through NeuroLink with service-account authentication, Gemini 2.5 extended context, multimodal inputs, structured output, and retry handling. By the end of this tutorial, you will know when caller-supplied regions apply, how to choose stable Gemini tiers, and how to invoke Claude models hosted on Vertex.
Configuring Vertex AI with NeuroLink
Vertex AI authentication differs from API key-based providers. Google Cloud uses service accounts and application default credentials for secure access.
Authentication Options
NeuroLink uses environment variables for Vertex AI authentication. Configure your credentials before using the SDK:
1
2
3
export GOOGLE_CLOUD_PROJECT="your-gcp-project"
export GOOGLE_APPLICATION_CREDENTIALS="/path/to/service-account.json"
export GOOGLE_CLOUD_LOCATION="us-central1"
Then use the NeuroLink SDK:
1
2
3
4
5
6
7
8
9
10
11
12
13
import { NeuroLink, VertexModels } from '@juspay/neurolink';
const neurolink = new NeuroLink();
// Basic generation with Vertex AI
const response = await neurolink.generate({
input: { text: "Explain machine learning concepts" },
provider: "vertex",
model: VertexModels.GEMINI_2_5_PRO, // "gemini-2.5-pro"
maxTokens: 1024
});
console.log(response.content);
You can also specify a region per-request for non-Gemini models on Vertex, such as Claude:
1
2
3
4
5
6
const response = await neurolink.generate({
input: { text: "Process this request in EU region" },
provider: "vertex",
model: "claude-sonnet-4-6",
region: "europe-west4" // Honored for Claude-on-Vertex; ignored for Gemini
});
Gemini models are always routed to Vertex’s global endpoint regardless of any region you pass – the region option only takes effect for non-Gemini models such as Claude on Vertex.
Required IAM Permissions
Your service account needs these roles:
| Role | Purpose |
|---|---|
roles/aiplatform.user | Invoke Vertex AI endpoints |
roles/storage.objectViewer | Access GCS for multimodal inputs |
roles/bigquery.dataViewer | Query BigQuery for context (optional) |
1
2
3
4
5
6
7
8
# Grant required roles
gcloud projects add-iam-policy-binding PROJECT_ID \
--member="serviceAccount:neurolink-sa@PROJECT_ID.iam.gserviceaccount.com" \
--role="roles/aiplatform.user"
gcloud projects add-iam-policy-binding PROJECT_ID \
--member="serviceAccount:neurolink-sa@PROJECT_ID.iam.gserviceaccount.com" \
--role="roles/storage.objectViewer"
Regional Endpoints
Vertex AI operates regionally, but NeuroLink routes every Gemini model (gemini-*) to Vertex’s global endpoint regardless of the region you pass – the caller-supplied region is honored only for non-Gemini Vertex models, such as Claude. For data residency requirements on Gemini, set the project’s default location instead:
1
2
3
4
5
6
7
8
9
10
11
import { NeuroLink, VertexModels } from '@juspay/neurolink';
const neurolink = new NeuroLink();
// GOOGLE_CLOUD_LOCATION / GOOGLE_VERTEX_LOCATION sets the project default;
// Gemini requests still resolve to the global endpoint internally.
const response = await neurolink.generate({
input: { text: "Analyze this document" },
provider: "vertex",
model: VertexModels.GEMINI_2_5_PRO,
});
Gemini Model Capabilities
NeuroLink’s Vertex model catalog includes stable Gemini 2.5 tiers as well as newer preview entries. The examples below use stable Gemini 2.5 model IDs for long-lived application code.
Model Selection
Choose the right model for your use case:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
import { NeuroLink, VertexModels } from '@juspay/neurolink';
const neurolink = new NeuroLink();
// Gemini 2.5 Flash - Fast responses, cost-effective
const flashResponse = await neurolink.generate({
input: { text: "Summarize this article in one sentence" },
provider: "vertex",
model: VertexModels.GEMINI_2_5_FLASH, // "gemini-2.5-flash"
maxTokens: 100
});
// Gemini 2.5 Pro - Balanced performance
const proResponse = await neurolink.generate({
input: { text: "Analyze the quarterly financial report and identify key trends" },
provider: "vertex",
model: VertexModels.GEMINI_2_5_PRO, // "gemini-2.5-pro"
maxTokens: 2000
});
// Gemini 2.5 Pro - Maximum capability for complex tasks
const complexResponse = await neurolink.generate({
input: { text: "Develop a comprehensive market entry strategy for..." },
provider: "vertex",
model: VertexModels.GEMINI_2_5_PRO, // "gemini-2.5-pro"
maxTokens: 8000
});
Extended Context Windows
Gemini 2.5 supports context windows up to 1 million tokens. NeuroLink passes your content efficiently:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
import { NeuroLink, VertexModels } from '@juspay/neurolink';
const neurolink = new NeuroLink();
// Combine multiple documents into context
const documentContext = documents.map(doc => doc.text).join('\n\n---\n\n');
// Generate with large context
const response = await neurolink.generate({
input: {
text: `${documentContext}\n\nSummarize the key legal arguments from these documents.`
},
provider: "vertex",
model: VertexModels.GEMINI_2_5_PRO,
systemPrompt: "You are a legal analyst with access to the complete case file.",
maxTokens: 4096
});
console.log(response.content);
Multimodal Processing
Gemini models excel at understanding multiple modalities. NeuroLink provides a consistent interface for images, video, and documents.
Image Understanding
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
import { NeuroLink, VertexModels } from '@juspay/neurolink';
import * as fs from 'fs';
const neurolink = new NeuroLink();
// Analyze image from local file
const imageBuffer = fs.readFileSync('product-photo.jpg');
const response = await neurolink.generate({
input: {
text: "Describe this image in detail. What objects, people, and activities are visible?",
images: [imageBuffer]
},
provider: "vertex",
model: VertexModels.GEMINI_2_5_PRO,
maxTokens: 1024
});
console.log(response.content);
// Analyze image from URL
const urlResponse = await neurolink.generate({
input: {
text: "Extract all text visible in this image",
images: ["https://storage.googleapis.com/bucket/image.jpg"]
},
provider: "vertex",
model: VertexModels.GEMINI_2_5_PRO
});
// Multiple images comparison
const beforeImage = fs.readFileSync('before.jpg');
const afterImage = fs.readFileSync('after.jpg');
const comparisonResponse = await neurolink.generate({
input: {
text: "Compare these two images. What changed?",
images: [beforeImage, afterImage]
},
provider: "vertex",
model: VertexModels.GEMINI_2_5_PRO
});
Video Analysis
Process video files with Gemini:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
import { NeuroLink, VertexModels } from '@juspay/neurolink';
import * as fs from 'fs';
const neurolink = new NeuroLink();
const videoBuffer = fs.readFileSync('presentation.mp4');
// Analyze video content
const videoResponse = await neurolink.generate({
input: {
text: "Create a detailed transcript with timestamps. Identify speakers if possible.",
videoFiles: [videoBuffer]
},
provider: "vertex",
model: VertexModels.GEMINI_2_5_PRO,
maxTokens: 4096
});
console.log(videoResponse.content);
Document Understanding
Process PDFs directly:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
import { NeuroLink, VertexModels } from '@juspay/neurolink';
import * as fs from 'fs';
const neurolink = new NeuroLink();
// Analyze PDF document
const pdfBuffer = fs.readFileSync('contract.pdf');
const pdfResponse = await neurolink.generate({
input: {
text: "Extract key terms from this contract: parties, effective date, payment terms, and termination clauses.",
pdfFiles: [pdfBuffer]
},
provider: "vertex",
model: VertexModels.GEMINI_2_5_PRO,
maxTokens: 2048
});
// Multi-page document processing
const reportBuffer = fs.readFileSync('annual-report.pdf');
const reportAnalysis = await neurolink.generate({
input: {
text: "Create an executive summary. Include key metrics, YoY changes, and strategic initiatives.",
pdfFiles: [reportBuffer]
},
provider: "vertex",
model: VertexModels.GEMINI_2_5_PRO,
maxTokens: 4096
});
Extended Thinking Capabilities
Gemini 2.5 models support advanced reasoning for complex tasks. Pass thinkingConfig with a token budget and NeuroLink translates it into the wire shape Gemini 2.5 expects:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
import { NeuroLink, VertexModels } from '@juspay/neurolink';
const neurolink = new NeuroLink();
// Enable extended thinking with a token budget for complex analysis
const response = await neurolink.generate({
input: { text: "Solve this complex optimization problem step by step..." },
provider: "vertex",
model: VertexModels.GEMINI_2_5_PRO,
thinkingConfig: {
enabled: true,
budgetTokens: 8000
},
maxTokens: 8000
});
console.log(response.content);
// Use appropriate models based on task complexity
const quickAnalysis = await neurolink.generate({
input: { text: "Summarize this article" },
provider: "vertex",
model: VertexModels.GEMINI_2_5_FLASH,
maxTokens: 1024
});
Structured Output with Schemas
Get type-safe JSON responses using Zod schemas:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
import { NeuroLink, VertexModels } from '@juspay/neurolink';
import { z } from 'zod';
const neurolink = new NeuroLink();
// Define schema for structured output
const ProductAnalysis = z.object({
productName: z.string(),
features: z.array(z.string()),
pros: z.array(z.string()),
cons: z.array(z.string()),
rating: z.number().min(1).max(5)
});
const response = await neurolink.generate({
input: { text: "Analyze the iPhone 16 Pro Max" },
provider: "vertex",
model: VertexModels.GEMINI_2_5_PRO,
schema: ProductAnalysis
});
console.log(response.structuredData); // Parsed object matching the schema
Enterprise Deployment Patterns
Production Vertex AI deployments require attention to security and cost management.
Environment Configuration
Configure Vertex AI through environment variables:
1
2
3
4
# Required configuration
export GOOGLE_CLOUD_PROJECT="your-project"
export GOOGLE_APPLICATION_CREDENTIALS="/path/to/service-account.json"
export GOOGLE_CLOUD_LOCATION="us-central1"
Cost-Effective Model Selection
Choose appropriate models based on task complexity:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
import { NeuroLink, VertexModels } from '@juspay/neurolink';
const neurolink = new NeuroLink();
// Use Flash for simple tasks (lower cost)
async function quickTask(prompt: string) {
return neurolink.generate({
input: { text: prompt },
provider: "vertex",
model: VertexModels.GEMINI_2_5_FLASH,
maxTokens: 500
});
}
// Use Pro for complex analysis
async function complexAnalysis(prompt: string) {
return neurolink.generate({
input: { text: prompt },
provider: "vertex",
model: VertexModels.GEMINI_2_5_PRO,
maxTokens: 4096
});
}
// Use Flash Lite for high-volume, simple tasks
async function bulkProcessing(prompts: string[]) {
return Promise.all(
prompts.map(prompt =>
neurolink.generate({
input: { text: prompt },
provider: "vertex",
model: VertexModels.GEMINI_2_5_FLASH_LITE,
maxTokens: 256
})
)
);
}
Error Handling and Retries
Implement robust error handling:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
import { NeuroLink, VertexModels } from '@juspay/neurolink';
const neurolink = new NeuroLink();
async function generateWithRetry(
prompt: string,
maxRetries = 3
): Promise<string> {
for (let attempt = 1; attempt <= maxRetries; attempt++) {
try {
const response = await neurolink.generate({
input: { text: prompt },
provider: "vertex",
model: VertexModels.GEMINI_2_5_PRO,
timeout: 30000 // 30 second timeout
});
return response.content;
} catch (error) {
if (attempt === maxRetries) throw error;
// Exponential backoff
await new Promise(resolve =>
setTimeout(resolve, Math.pow(2, attempt) * 1000)
);
}
}
throw new Error('Max retries exceeded');
}
Function Calling with Tools
Gemini models support function calling:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
import { NeuroLink, VertexModels, tool } from '@juspay/neurolink';
import { z } from 'zod';
const neurolink = new NeuroLink();
// Define tools using Vercel AI SDK format
const weatherTool = tool({
description: 'Get current weather for a location',
parameters: z.object({
location: z.string().describe('City name'),
units: z.enum(['celsius', 'fahrenheit']).optional()
}),
execute: async ({ location, units }) => {
// Implement weather lookup
return { temperature: 22, condition: 'sunny', location };
}
});
const response = await neurolink.generate({
input: { text: "What's the weather in Tokyo?" },
provider: "vertex",
model: VertexModels.GEMINI_2_5_PRO,
tools: { weather: weatherTool }
});
console.log(response.content);
if (response.toolCalls) {
console.log('Tool calls:', response.toolCalls);
}
Claude Models via Vertex AI
Access Anthropic Claude models through Vertex AI:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
import { NeuroLink, VertexModels } from '@juspay/neurolink';
const neurolink = new NeuroLink();
// Claude 4.6 Sonnet on Vertex AI
const claudeResponse = await neurolink.generate({
input: { text: "Explain the benefits of using Claude through Vertex AI" },
provider: "vertex",
model: VertexModels.CLAUDE_4_6_SONNET, // "claude-sonnet-4-6"
maxTokens: 2048
});
// Claude for complex reasoning tasks
const reasoningResponse = await neurolink.generate({
input: { text: "Solve this complex reasoning problem..." },
provider: "vertex",
model: VertexModels.CLAUDE_4_6_SONNET, // "claude-sonnet-4-6"
maxTokens: 4096
});
console.log(reasoningResponse.content);
Note: Model names and IDs in code examples reflect versions available at time of writing. Model availability, naming conventions, and pricing change frequently. Always verify current model IDs with your provider’s documentation before deploying to production.
What You Built
You configured Vertex AI authentication with service accounts, selected the right Gemini model for each task (Flash for speed, Pro for reasoning), processed multimodal inputs including images, video, and PDF documents through a consistent interface, used extended thinking and structured output for complex analysis, and added retry logic and function calling for production reliability.
Note: NeuroLink’s Vertex catalog includes
gemini-3.1-pro-preview, but preview IDs can change. The examples use stable Gemini 2.5 models; check the Vertex AI model documentation before choosing a preview for production.
Start with simple text generation, then progressively add multimodal inputs as your use case demands. The combination of Gemini’s capabilities and NeuroLink’s unified interface creates powerful AI applications on Google Cloud infrastructure.
Continue with these related tutorials:
- Building AI Discord Bots with NeuroLink – Automatic provider switching when Vertex is unavailable
- Structured Output – Type-safe JSON responses from Gemini models
- Streaming Patterns – Real-time response delivery
- Cost Optimization – Reduce AI spending across providers
Related posts:
