Scientyfic World

AWS Amplify and AWS CDK: How to Connect MySQL and PostgreSQL Databases?

If you already have a MySQL or PostgreSQL database (on Amazon RDS, Aurora or similar) and want a GraphQL API in front of it, AWS Amplify can generate one for...

Share:

Get an AI summary of this article

Connect PostgreSQL MySQL Databases using AWS Amplify and AWS CDK

If you already have a MySQL or PostgreSQL database (on Amazon RDS, Aurora or similar) and want a GraphQL API in front of it, AWS Amplify can generate one for you: it turns your SQL tables into a GraphQL API served by AWS AppSync, with a Lambda function that runs the queries. You can set this up in two ways: with Amplify’s Gen 2 tooling (simplest), or with the AWS CDK and Amplify’s GraphQL API construct (more control, supports private databases). This guide covers both, with the CDK version in detail.

Updated October 2026: I rewrote this January 2024 guide. Amplify now has a newer “Gen 2” way of connecting existing SQL databases, which the original didn’t cover; the CDK construct’s VPC settings use a different shape than the old post showed (subnetAvailabilityZoneConfig entries rather than a flat list of subnet IDs); and the old post had no tested code. I wrote the CDK stack below against the current @aws-amplify/graphql-api-construct package (version 1.22.3, published September 2026), type-checked it with TypeScript in strict mode, and ran cdk synth-equivalent synthesis offline to confirm it produces an AppSync API, a GraphQL schema, an API key, a data source and a Lambda function. I did not deploy it to AWS or run it against a real database, so test in a throwaway account first.

How It Works

  1. A client sends a GraphQL query or mutation to an AWS AppSync API.
  2. AppSync passes the request to a Lambda function that Amplify generates (the “SQL Lambda”).
  3. The Lambda reads the database connection details from Systems Manager Parameter Store (or Secrets Manager), connects to your database, runs SQL built from your schema, and returns the rows.

Your GraphQL types map to tables: the schema you write describes the table’s columns, and Amplify generates create, read, update, delete and list operations. For queries the generator can’t express, you write your own SQL and attach it to a GraphQL field with the @sql directive.

Choose Your Approach

Amplify Gen 2 (ampx CLI)AWS CDK with the GraphQL API construct
Best forFast start, a full-stack Amplify appInfrastructure-as-code, existing CDK projects, private databases
Database connectionOne connection-string secretSeparate SSM parameters (host, port, name, user, password) or a Secrets Manager secret
NetworkPer Amplify’s docs, a database in a VPC must be configured as publicly accessible, with its security group allowing the database port and HTTPSCan place the Lambda inside your VPC, so the database can stay private
SchemaGenerated from the live database with a commandYou write the GraphQL schema
ControlLessMore

The security difference matters. A publicly accessible database is an internet-facing attack surface, even with a restrictive security group. For anything beyond a prototype, prefer the CDK route with the database in private subnets.

Option A: The Quick Route With Amplify Gen 2

This follows Amplify’s documentation for connecting an existing Postgres or MySQL database. From your Amplify Gen 2 project:

# 1. Store the connection string as a secret (mysql://user:password@host:port/db or postgres://...)
npx ampx sandbox secret set SQL_CONNECTION_STRING

# 2. Generate a schema file from the database (only tables with a primary key are included)
npx ampx generate schema-from-database \
  --connection-uri-secret SQL_CONNECTION_STRING \
  --out amplify/data/schema.sql.ts

Then combine the generated schema with your other data models in amplify/data/resource.ts and deploy:

import { schema as generatedSqlSchema } from './schema.sql';

const sqlSchema = generatedSqlSchema.authorization(allow => allow.guest());
const combinedSchema = a.combine([schema, sqlSchema]);
npx ampx sandbox

The documentation also recommends RDS Proxy when concurrent queries could exceed the database’s connection limit, and says production deployments need the same secret name added in Amplify Hosting. Note the allow.guest() rule above lets unauthenticated users access the data; change it to a real authorization rule before exposing anything real. Remember the caveat above about public accessibility.

Option B: The CDK Route in Detail

Prerequisites

  • An AWS account, the AWS CLI configured with credentials, and a Region.
  • Node.js and npm. The CDK CLI package (aws-cdk) currently declares Node.js 18 or later, but use a current LTS release.
  • The CDK CLI, installed as in the CDK getting-started guide (npm install -g aws-cdk), and a one-time cdk bootstrap in each account and Region.
  • A MySQL or PostgreSQL database that your tables already exist in, and its VPC, subnet and security-group IDs if it is in a VPC.
  • Each table you want exposed needs a primary key, which Amplify requires to generate operations.

Step 1: Store the Database Connection Details

The construct reads five values from Systems Manager Parameter Store: hostname, port, database name, username and password. Create them as SecureString parameters (encrypted with KMS) in the same Region as the stack:

aws ssm put-parameter --name /app/db/hostname --type SecureString --value "mydb.abc123.us-east-1.rds.amazonaws.com"
aws ssm put-parameter --name /app/db/port     --type SecureString --value "5432"
aws ssm put-parameter --name /app/db/name     --type SecureString --value "appdb"
aws ssm put-parameter --name /app/db/username --type SecureString --value "app_user"
aws ssm put-parameter --name /app/db/password --type SecureString --value "REPLACE_WITH_A_STRONG_PASSWORD"

Don’t reuse an administrator account here. Create a database user with only the privileges the API needs (for example read and write on specific tables). The construct also supports pulling credentials from an AWS Secrets Manager secret instead, which is the better option if you want automatic password rotation. Avoid typing real passwords on a shared shell; they land in shell history, so use your shell’s history controls or the console.

Step 2: Create the CDK Project

mkdir sql-graphql-api && cd sql-graphql-api
npx aws-cdk init app --language typescript
npm install @aws-amplify/graphql-api-construct

The construct package requires aws-cdk-lib 2.224 or newer and constructs 10.5 or newer as peer dependencies; the project template installs compatible versions.

Step 3: Define the Schema and the Stack

Replace the stack code (in lib/, or the app file in bin/ for a quick test) with the following. The GraphQL schema lives in a string; each SQL-backed type uses @model as usual, @refersTo to map a GraphQL type to a differently named table, and @primaryKey to mark the key. The query at the bottom runs custom SQL, referenced by name from customSqlStatements:

import { App, CfnOutput, Duration, Stack } from 'aws-cdk-lib';
import {
  AmplifyGraphqlApi,
  AmplifyGraphqlDefinition,
  SQLLambdaModelDataSourceStrategy,
} from '@aws-amplify/graphql-api-construct';

const app = new App();
const stack = new Stack(app, 'SqlGraphqlStack', { env: { account: '111122223333', region: 'us-east-1' } });

// Where the SQL-resolving Lambda finds the database. These are SecureString parameters
// in Systems Manager Parameter Store that you create yourself (Step 1).
const dataSourceStrategy: SQLLambdaModelDataSourceStrategy = {
  name: 'AppDatabase',
  dbType: 'POSTGRES',                        // or 'MYSQL'
  dbConnectionConfig: {
    hostnameSsmPath: '/app/db/hostname',
    portSsmPath: '/app/db/port',
    databaseNameSsmPath: '/app/db/name',
    usernameSsmPath: '/app/db/username',
    passwordSsmPath: '/app/db/password',
  },
  // Needed when the database is in a VPC, e.g. an RDS instance or proxy
  vpcConfiguration: {
    vpcId: 'vpc-0123456789abcdef0',
    securityGroupIds: ['sg-0123456789abcdef0'],
    subnetAvailabilityZoneConfig: [
      { subnetId: 'subnet-0aaaaaaaaaaaaaaaa', availabilityZone: 'us-east-1a' },
      { subnetId: 'subnet-0bbbbbbbbbbbbbbbb', availabilityZone: 'us-east-1b' },
    ],
  },
  // Custom SQL referenced by name from the schema with @sql(reference: "...")
  customSqlStatements: {
    getUserByEmail: 'SELECT * FROM users WHERE email = :email',
  },
};

const api = new AmplifyGraphqlApi(stack, 'SqlApi', {
  apiName: 'AppApi',
  definition: AmplifyGraphqlDefinition.fromString(/* GraphQL */ `
    type User @model @refersTo(name: "users") {
      id: ID! @primaryKey
      name: String!
      email: String
    }

    type Query {
      getUserByEmail(email: String!): User @sql(reference: "getUserByEmail")
    }
  `, dataSourceStrategy),
  authorizationModes: {
    apiKeyConfig: { expires: Duration.days(30) },   // demo only: use Cognito or IAM for real apps
  },
});

new CfnOutput(stack, 'GraphqlUrl', { value: api.graphqlUrl });

I compiled this with TypeScript 5 in strict mode and synthesized it offline: it produced the AppSync API, schema, API key, data source and Lambda function. Notes on each part:

  • dbType is 'POSTGRES' or 'MYSQL'.
  • dbConnectionConfig lists the five SSM parameter paths from Step 1.
  • vpcConfiguration places the SQL Lambda inside your VPC so it can reach a private database. It takes the VPC ID, security group IDs, and one subnet per Availability Zone as { subnetId, availabilityZone } pairs. Per the construct’s documentation, it also creates the VPC service endpoints and inbound rules on port 443 that the Lambda needs to read the SSM parameters. Your database’s security group must allow inbound connections on the database port (5432 or 3306) from the Lambda’s security group. Remove this block if your database is publicly reachable (not recommended).
  • customSqlStatements holds the SQL for @sql(reference: "name") fields. Use named parameters (:email) rather than building strings, so inputs are bound as parameters, which protects against SQL injection. The original post’s inline @sql(statement: ...) style still exists, but keeping SQL in customSqlStatements (or SQL files) is easier to review and test.
  • authorizationModes: the API key above is for a demo only; the construct creates one that expires after the number of days you set (here 30). For real applications use a Cognito user pool, IAM or OIDC and add @auth rules to your types so only the right users can read or change data.
  • The stack’s env uses a dummy account ID so the test could run offline. Replace it with your own, or remove env to use your CLI’s defaults.

Step 4: Deploy and Test

npx cdk diff
npx cdk deploy

cdk diff shows what will be created, and it is worth reading, since this stack creates IAM roles, a Lambda function, VPC endpoints and an AppSync API. After deployment, the stack output GraphqlUrl (added in the code above) gives the endpoint. Fetch the API key from the AppSync console (or add it to the outputs, bearing in mind that outputs are visible to anyone who can read the stack). Then query it:

curl -s -X POST "$GRAPHQL_URL" \
  -H "Content-Type: application/json" \
  -H "x-api-key: $API_KEY" \
  -d '{"query":"query { listUsers { items { id name email } } }"}'

curl -s -X POST "$GRAPHQL_URL" \
  -H "Content-Type: application/json" \
  -H "x-api-key: $API_KEY" \
  -d '{"query":"query { getUserByEmail(email: \"[email protected]\") { id name } }"}'

The exact shape of the generated list operation (for example listUsers with an items list) depends on the models in your schema, so check the generated schema in the AppSync console. When you are done experimenting, npx cdk destroy removes the stack.

Going to Production

  • Keep the database private and use the VPC configuration. Use security groups that allow only the SQL Lambda to connect.
  • Use RDS Proxy (AWS documentation) so a burst of GraphQL requests doesn’t exhaust the database’s connections. Lambda scales out quickly and each concurrent execution can hold a connection.
  • Authorization first. Replace the API key with Cognito, IAM or OIDC, write @auth rules, and test them, because a SQL-backed API is a direct window onto your data.
  • Least-privilege database user, rotated credentials (use Secrets Manager rotation), and no administrator accounts in the API path.
  • Expect Lambda cold starts and a VPC. The construct has an option for provisioned concurrency on the SQL Lambda if latency matters; it adds cost.
  • Watch the GraphQL cost model. Deeply nested queries can fan out into many SQL queries (the N+1 problem). Limit depth, paginate list queries, and add indexes for columns you filter on.
  • Schema drift. Your GraphQL schema describes the tables, so a database migration that renames or removes a column can break the API. Run migrations and API changes through the same review process, and test against a staging database.
  • Monitor. Use CloudWatch logs and alarms for the Lambda and AppSync errors and latency.

For broader context on putting an app on AWS, see how to deploy a web app on AWS and our serverless architecture guide.

Troubleshooting

SymptomWhat to check
The SQL Lambda times out connecting to the databaseSecurity group rules: the database’s group must allow the Lambda’s group on the database port. Subnet routing, and that the Lambda is in the same VPC as the database.
The Lambda can’t read the SSM parametersParameter names match exactly, they exist in the same Region, and the VPC endpoint rules (port 443) are in place. Check CloudWatch logs for access-denied errors.
A table is missing from the generated schema (Gen 2)It probably has no primary key; Amplify only generates models for tables with a primary key.
“Too many connections” errorsPut RDS Proxy in front of the database and reduce Lambda concurrency.
Authorization errors from AppSyncThe authorization mode you call with must be configured in authorizationModes and allowed by the type’s @auth rules; API keys expire.
Deploy fails on VPC settingsSubnets must be in the stated VPC, one per Availability Zone, and the security groups must be in the same VPC.

Can AWS Amplify connect to an existing MySQL or PostgreSQL database?

Yes. Amplify can generate a GraphQL API on AWS AppSync for existing MySQL and PostgreSQL databases, through Amplify Gen 2 (using the ampx CLI and a connection-string secret) or through the AWS CDK GraphQL API construct with a SQL data source strategy.

Does my database have to be publicly accessible?

With the Gen 2 route, Amplify’s documentation says a database in a VPC must be set to publicly accessible. With the CDK construct you can configure the SQL Lambda to run inside your VPC, so the database can stay private, which is the better choice for production.

Where should I store the database credentials?

In AWS Systems Manager Parameter Store as SecureString parameters, or in AWS Secrets Manager, which also supports rotation. Never put them in your schema, code or repository.

Why is a table missing from my generated schema?

Amplify generates models only for tables that have a primary key. Add a primary key to the table, or define the model manually.

How do I avoid too many database connections?

Use Amazon RDS Proxy between the SQL Lambda and the database, and consider limiting the Lambda’s concurrency. Lambda can scale out quickly, and each execution environment may open its own connection.

Snehasish Konger
Developed @scientyficworld.org | Technical writer @Nected | Content Developer
Connect with Snehasish Konger

On This page

Take a Pause with Intervals

A Sunday letter on building, writing, and thinking deeper as a developer — short, honest, and worth your time.

Snehasish Konger profile photo

"Hey there — I'm Snehasish. Hope this post saved you some head-scratching time! I've spent years turning technical chaos into clarity, and I'm here to be your guide through the maze of modern tech. Stick around for more lightbulb moments — we're just getting started."

Related Posts