OMX Helsinki — S&P 500 — DAX — NASDAQ 100 — STOXX 600 — EUR/USD — EUR/SEK — BTC/USD — ETH/USD — Euribor 3M — Euribor 12M —

Autonomous agents and skills in practice — how a startup can put them to work

The big shift with agents: AI no longer just answers, it starts doing. A skill is a specialized capability you give an agent — and it's worth building one when you notice yourself pasting the same instructions into an AI over and over.

Here's the big shift with agents: AI no longer just answers, it starts doing.

That matters for startups, because a small team can get access to “digital coworkers” that handle code, documents, research, spreadsheets, customer work, and recurring operations. But that's exactly why the boundaries need to be defined precisely.

1. What does a skill mean in practice?

A skill is a specialized capability you give an agent. For example, it could be:

  • “analyze the git changes before committing”
  • “summarize the customer interviews”
  • “turn the CSV report into a table for the leadership team”
  • “find the bug based on the tests”
  • “draft a sales email but don't send it”
  • “update the documentation to reflect the new feature”

According to the Claude Code documentation, a skill is built, for example, as a file named SKILL.md, which gives instructions on how Claude should use that skill. A skill can be invoked directly with a command such as /skill-name, or Claude can use it on its own when it recognizes that a task is a good fit.[1]

An important practical point: build a skill when you notice yourself pasting the same instructions into an AI over and over.

2. Claude Code in Visual Studio Code: getting started in practice

Claude Code runs in VS Code as its own extension. Anthropic recommends the VS Code extension as the native way to use Claude Code in the editor. It brings features such as inline diffs, file @-mentions, plan review, and keyboard shortcuts.[2]

In practice, a startup's development team could do this:

Step 1: Install Claude Code in VS Code

Open the VS Code Extensions view and search for “Claude Code”. According to the documentation, you need VS Code 1.98.0 or later and an Anthropic account.[2]

Step 2: Open the project and sign in

Once the extension is installed, Claude Code appears in VS Code as a Spark icon. The first time you use it, you sign in through your browser.[2]

Step 3: Add a CLAUDE.md file to the project

A startup should create an instructions file in the project root that states:

# Project instructions for Claude

• Use TypeScript.
• Don't change the database schema without permission.
• Don't make commits without human approval.
• Always write tests alongside any new feature.
• Follow the existing architecture.
• If you're unsure, propose a plan before making changes.

These are the agent's “workplace ground rules.”

Step 4: Create your first skill

For example, a startup could build a skill that reviews pending changes before a pull request. The folder structure could look like this:

.claude/
  skills/
    review-before-pr/
      SKILL.md

SKILL.md could look like this:

---
name: review-before-pr
description: Reviews git changes before a pull request and looks for bugs, missing tests, and risks.
---

Act as a senior software engineer.

Follow these steps:

1) Look at the current git changes.
2) Identify potential bugs.
3) Check for missing tests.
4) Identify security risks.
5) Suggest improvements.

Don't make commits. Don't change files without separate permission.

Finish with a prioritized list: critical fixes, recommended fixes, minor improvements.

Usage: /review-before-pr

This is a practical skill because it saves senior developers' time and raises quality before the code goes into PR review.

3. Codex: how does an autonomous coding agent work?

OpenAI's Codex is a cloud-based software engineering agent. OpenAI describes it as an agent that can write features, answer questions about the codebase, fix bugs, and propose pull requests. Each task runs in its own isolated cloud environment preloaded with the repository.[3]

In practice, here's how it differs from an ordinary chat-based coding assistant:

  • an ordinary chat gives you an answer
  • Codex gets a task
  • Codex reads the codebase
  • Codex makes changes
  • Codex runs tests, linters, and type checks
  • Codex can propose a PR

According to OpenAI, Codex can read and edit files and run commands such as tests, linters, and type checkers.[3]

A good Codex task for a startup might be:

Fix the checkout flow bug where the discount code disappears if the user changes the shipping method.

Boundaries:

• Don't change the payment provider integration.
• Add a regression test.
• Run the tests.
• Finish with a summary of the changes and risks.

A bad task would be: “Improve checkout.” Too open-ended. Too risky.

4. Claude Cowork: an agent for knowledge work, not just code

Claude Cowork is Anthropic's agentic workspace for knowledge work. Anthropic describes it as a system you give a goal to, and Claude then works on your computer, in your local files and applications, to produce a finished result.[4]

It's especially well suited to a startup's non-technical work:

  • pulling together investor materials
  • analyzing customer feedback
  • organizing documents
  • condensing research material
  • extracting contract clauses
  • compiling recruiting feedback
  • preparing sales materials

Anthropic says Claude Cowork is built for situations where chat isn't enough and the user wants to hand off an outcome rather than steer every intermediate step themselves.[4]

Startup example:

Go through the folder /customer-interviews.
Pull out every mention of pricing, onboarding, and integrations.
Put them in a table: customer, observation, pain point, possible product improvement, urgency.

Don't delete or move any files.
Finish with a one-page summary for the founding team.

This is a good agent task because it's well scoped, recurring, and produces material for decision-making.

5. OpenClaw: an open-source personal agent

OpenClaw is an open-source personal AI assistant that you run on your own devices. According to its GitHub description, it works in the channels you already use and supports WhatsApp, Telegram, Slack, Discord, Google Chat, Signal/iMessage, and Microsoft Teams, among others.[5]

What makes OpenClaw interesting from a startup perspective is this: it shows where the user interface for agents may be heading. An agent doesn't necessarily live in a single app. It can be present in the channels where the team already works.

A startup could try an OpenClaw-style model like this, for example:

  • a Telegram bot for the founders
  • a Slack agent for internal team questions
  • a WhatsApp agent for personal reminders
  • a locally run assistant that doesn't require sending all your data to the cloud

But you need to be especially careful here: if the agent can access messages, calendars, files, or customer data, its permissions have to be tightly restricted.

6. A practical rollout plan for startups

Don't kick off your agent strategy by saying: “Let's automate everything.”

1. Pick one high-friction process

Good first candidates:

  • initial bug triage
  • PR review before a human looks at it
  • analyzing customer feedback
  • sales lead enrichment
  • drafting investor updates
  • writing release notes
  • updating onboarding documentation

2. Decide on the agent's permissions

Use three levels:

Can do on its own: read files, draft content, run tests, suggest changes, write summaries.

Can suggest: code changes, CRM updates, customer messages, roadmap changes.

Requires permission: commits, merging pull requests, sending email, changing customer data, approving invoices, changing the production environment.

3. Build your first skills

A startup's first five useful skills could be:

  • /review-before-pr — Reviews code changes before a PR.
  • /write-tests — Suggests and writes tests for a new feature.
  • /customer-feedback-summary — Condenses customer feedback into product insights.
  • /investor-update — Drafts a monthly investor update from the metrics you provide.
  • /release-notes — Turns git changes into release notes customers can understand.

4. Measure the benefit

Don't just measure whether it “feels handy.” Measure:

  • time saved
  • fewer bugs
  • PR cycle time
  • how up to date the documentation is
  • how quickly customer feedback gets processed
  • number of errors made by the agent
  • how much human correction is needed
  • cost per task

7. A concrete startup example: a SaaS company rolls out agents

Picture an 8-person B2B SaaS startup. The team: 3 developers, 1 designer, 1 salesperson, 1 customer support person, 1 founder-CEO, 1 growth/marketing person.

Week 1: Claude Code in VS Code

The developers start using Claude Code. The initial boundaries:

Claude may: read code, suggest refactoring, write tests, make small changes with approval.

Claude may not: change the database schema, deploy to production, delete files, make commits without permission.

Week 2: The first skills

Create /review-before-pr, /write-tests, /explain-module. Goal: junior and senior developers both save time, but a human approves everything.

Week 3: Codex for bug fixes

Codex is given well-scoped issues:

Fix this bug. Add a test. Don't change the public API. Run the tests. Write a summary.

Good use: small bugs in parallel. Bad use: “build a new payment system from start to finish.”

Week 4: Claude Cowork for customer feedback

Cowork is given a folder of customer interviews. The task: the most common pain points, the top feature requests, pricing comments, and 5 recommendations for the roadmap — without modifying the original files.

Week 5: Trying OpenClaw as an internal assistant

The team tests a Slack/Telegram-style agent that answers internal questions: “Where's the latest pitch deck?” “What did we agree on in the last customer meeting?” “When is the next release?”

Boundaries: no access to production data, no access to payroll information, no right to send external messages, read-only access to documents.

8. The best rule of thumb

When rolling out an agent, the most important question isn't: “What can this do?”

It's: “What should this be allowed to do without a human?”

A smart baseline for a startup is:

The agent can prepare. The agent can analyze. The agent can suggest.
A human approves.
The agent carries out only low-risk tasks.

9. A short template for instructing an agent

Copy this into every agent task:

Goal: [what you want to achieve]

Context: [why this is being done]

Allowed tools: [what the agent may use]

Can do on its own: [low-risk tasks]

May only suggest: [medium-risk tasks]

Requires permission: [high-risk tasks]

Output: [what kind of deliverable you want]

Finish by reporting: what you did, what you changed, what you didn't do, where you're unsure, and what you recommend next.

10. Conclusion

Startups shouldn't think of agents as “AI employees” who are given free rein. A better way to think about it: agents are fast specialist assistants, and you build skills for them precisely where the team is losing time.

  • Claude Code fits into the developer's editor.
  • Codex is suited to well-scoped coding tasks and fixing bugs in parallel.
  • Claude Cowork is suited to knowledge-work deliverables such as documents, research, and file handling.
  • OpenClaw shows what an open-source, channel-agnostic personal agent can look like.

A startup's competitive edge comes from using agents not as a flashy demo but as a quiet accelerator of everyday workflows.

Sources

  1. Extend Claude with skills — Claude Code Docs
  2. Use Claude Code in VS Code — Claude Code Docs
  3. Introducing Codex — OpenAI
  4. Claude Cowork — Anthropic's agentic AI for knowledge work
  5. OpenClaw — personal AI assistant on GitHub
Open the AI assistant chat. The chat loads only when you open it.