← Back to context

Comment by alin23

4 days ago

Lately I found MCP to be much more than a coding tool. For example, I implemented it in my more complex macOS apps [0] like rcmd, Clop, Lunar, so they can be configured by natural language.

So even with a local Qwen and Pi you can now say things like:

    Set up Clop to optimise any PNG that I drop in my website assets folder and convert to a webp with the same name near it

    Get Crank to start Time Machine backups immediately when I connect my HDD and notify me when the backup is done.

    I want to be able to hold rcmd and fuzzy search and focus cmux agent panes

BetterTouchTool has a great MCP which can create native SwiftUI views and bind them to hotkeys, trackpad gestures etc. It can leverage its immense macOS automation tools and private APIs to let agents do Computer Use.

You would need a much more capable coding model to code those tools from scratch and get the same fail-safe logic that the apps have honed over the years.

Like, since MCP, Crank [1] has fully replaced my use of crontab, launchd, scattered scripts I run once a week. Not that it could not do that before, but it's so much simpler now to just describe the automation and have it happen reliably and visible in the UI. The friction is gone.

[0] https://reddit.com/r/macapps/comments/1wkv0dy/mcp_in_macos_a...

[1] https://lowtechguys.com/crank

I gave claude an API key for Home Assistant. I can tell it to create dashboards, set up automations, diagnose problems etc, all using natural language. No MCP needed.

Yesterday I received a new thermometer for my aquarium to replace an old broken one. Both were bluetooth, but different models. I just told claude "I'm going to set up up my new bluetooth thermometer for my fish tank in a few minutes, keep an eye out for it and replace the old broken one with it in Home Assistant" and then walked away and put a battery in it and put it in my aquarium.

When I came back it had found it, replaced all my existing entities for the broken one with the new one, and verified it was all working with my existing graphs and automations.

  • MCP is a tool more for security than anything else. If you give your agents access to an API key, there's a chance they can accidentally or maliciously leak that key. If the MCP server has access to the keys instead, it takes that possibility away. That's not always something you need to care about, but it is very important for some people's threat model.

    • > MCP is a tool more for security than anything else. +1 to this; API keys mean: a) Your agents are overprivileged by default b) Your API key is more likely to get into logs and pre-training / leaked as FrinkleFrankie mentioned c) You can't manage it centrally with granular tool policies in Enterprises. (e.g. "you're not allowed to read slack channel for #finances")

      Also, you can always collapse MCP to CLI, so MCP as auth/security middleware makes a lot of sense.

      PoC of MCP -> CLI: https://github.com/edison-watch/cli

    • > MCP is a tool more for security than anything else.

      The agent can get to the resource through the MCP server or using API key. I personally do not see the benefit MCP is providing here. Sure you can reduce the exposed surface at MCP layer, but I do that at the API layer. I don't need to add another layer here.

      I can kinda understand if you do not have control of the API layer and/or you have to expose the API layer to the public Internet as well. Most of the time that is not the case for me.

      2 replies →

  • I +1 this approach. We are doing similarly. Instead of providing MCP, we are simply providing API key and documentation. Folks are able to then just paste that link and use our service.

    User will interact or build apps with simply text like: "Give me all the issues that are X, context: https://somedomain.com/llm.txt"

    llm.txt will have all the API instructions

  • One of the key points of the comment you replied to is that you don’t need a very powerful model to do MCP stuff, such as local Qwen. The “let the agent figure it out” approach works much better with powerful models, such as Claude (which you mentioned you are using).

  • Verging away from the topic, but I've found Claude is a great addition to HA. I want smart home features, but don't have the time/inclination to learn HA's way of working. Historically I just defaulted to Google Home because it was easier, but with Claude I barely need to touch HA configuration at all.

    Recently I wanted to set up a slightly complex routine involving some lights and a couple of motion sensors. It feels like magic to be able to describe the behavior I want, briefly discuss the implementation, and walk past the sensor and see it in action.

  • For Home Assistant, the MCP can enable the LLM to make changes without using up excessive tokens.

    For example, the only way to make a change to an HA automation via the API is to POST a whole new copy of the YAML, even if changing one things. The skill and HA MCP I use allows the LLM to use tools which make more precise changes without excessive context usage.

    Certainly still works either way though. And of course your LLM instance could just roll its own tools to do the exact same thing.

  • That's smart! That reminds me, I have to get back into HomeAssistant. I had my whole house through it a few years ago, but the complexity and things breaking in hard to debug ways made the experience too frustrating for my wife and visiting relatives.

    Having an agent keep an eye on stuff and fix things proactively should make the experience much better. Plus I can no longer write yamls at last.

    • Developing for HA with Claude has been great. Not only can it make all of those yaml changes based on natural language goals, but it's so easy now to create a custom dashboard or configure various apps. I was struggling with both the ChoreOps docs and its fairly cumbersome interface until I pointed Claude at it and told it what I was trying to do.

  • > No MCP needed

    Definitely helps that the home assistant api is documented online most likely in the training data.

    • One of the nice things about MCP is that the tools have descriptions that provide guidance to the LLM. You can provide documentation in other formats like OpenAPI but its nice that the documentation is so closely coupled with MCP servers.

      2 replies →

I just wanted to say thank you for making the tools that you have either free, or very reasonably priced. ZoomHider, MusicDecoy, YellowDot, and IsThereNet are some of the first things I install on my/my family's Macs (often before even Homebrew).

They're so powerful and yet get out of your way when you're not using them. I couldn't imagine being without them. Thanks y'all!

Hi! I've dabbled in implementing an MCP server/client back in March. To me a proper REST API and/or cli tool seems sufficient enough, agents use them with good efficiency. Any reason not to provide CLI or REST interface for your tools, but MCP for agents specifically?

  • In my apps, CLI came first which works through Mach ports as the IPC, so the "REST equivalent" of macOS apps is also present. MCP takes advantage of the same client-server architecture I created for the CLI, so there's nothing you can't do with the CLI that needs an MCP.

    But in my case, a CLI was not enough.

    Like, to the MCP I might say:

        Set Clop to make every video copied in ~/shots smaller, 2x and silent 
    

    Then the MCP can use elicitation and say:

        By smaller, you mean re-encode to compress file size (factor can be 0 to 100 max compression) or downscale resolution (100% same size, 50% half size)? 
        And does 2x mean faster speed? In which case do you want to keep frames so the video plays smoother or drop frames for size? Or does 2x mean upscale? 
    

    And the agent will present those as nice choice menus I can decide schematically on.

    With the CLI I have to first read, learn and memorize the requests and commands needed for each app, the accepted values and formats and the steps to reach a specific result.

    There's only so much space in my head I can leave for implementation details of arbitrary apps. I'd rather have an agent care about that.

    And yes I get the irony, those are my apps, I coded them by hand for years, I should know their implementation details, yet even I forget if I should pass 50% or 0.5 for half size.

    Btw Clop is a media file compressor for context: https://lowtechguys.com/clop

    • You answered "Why use MCP with your agent instead of using CLI manually?" but the more interesting question is "Why build an MCP when you could already point your agent to the CLI?". The user experience of "the agent will present those as nice choice menus I can decide schematically on" will probably be the same since agents are very adept at using cli tools and gathering required/optional arguments, examples, and warnings/errors to present you with useful choices on how to proceed.

      1 reply →

  • How do they authenticate to your API?? Do you want to ask normal people to store an API key and remember to rotate it every so often?

    • How does my browser authenticate to a website/service? Can't agents have session storage to store such information?

  • agents dont always have access to a terminal. why is this so difficult to imagine

    • Yes and harnesses can't provide any interface for an agent to send REST requests for example.

> Set up Clop to optimise any PNG that I drop in my website assets folder and convert to a webp with the same name near it

This seems like exactly the sort of thing I've done with shell scripts or even makefiles.

  • Nothing in there is innovative or unattainable, Clop uses open source tools behind the scenes anyway so obviously you can replicate it with scripts.

    But this is for people that already use the app, researchers, writers, students, people that aren't necessarily comfortable with a terminal. And given Clop already implements the basics: an efficient file events watcher, tuned encoders for the Mac silicon, fail safe backups and UI for seeing the result and interacting with it in real time, it has advantages over trying to do it yourself.

    • I'm not intending to denigrate your tool. I'm just commenting on the "look at how we go full circle" aspect.

      I will point out that the shell script way uses less resources than an LLM making a tool call. But I understand that these scenarios are not necessarily meant for the same user.

      2 replies →

So you found them useful to script what other open source tools can do with basic automation tools? How is any of this unique to MCP?

Why is AI involved for any other reason than building the original test implementation?

  • This is for giving the existing users of my apps the ability to describe what they want the app to do without having to navigate and learn the UI.

    Not sure if you got the right context, your question doesn't really make sense to me.

  • I find it exciting that the dust is finally settling.

    We're back onto the original use cases for natural language processing. This is where all the value always was, and now the market has proven to itself what anyone with even a bachelor's in computer science already knew.

Where I found MCP's really useful is integrating with "consumer AI" (chatgpt.com, claude.ai, etc).

I'm working on a sideproject called Rowbly[1]. It acts as a sharable data store where LLM's can dump research rather than keeping it in their memory or throwing it into a spreadsheet.

At first I thought "I don't need an MCP, I'll just expose a CLI" but that carries a pretty big limitation in that it only works with agents on your computer (Codex, Claude Code, Pi, etc). For the folks on here this is not an issue and is often times preferable but I'm also targeting the average LLM user that primarily interfaces with it via "consumer AI" and with those tools the stuff you can do is very very limited.

I still think MCP's have a long way to go maturity wise and hey maybe in a few years we will figure out a better way to do things but for now, if you want to interact with consumer AI apps, there's just no way around using them.

[1]: https://rowbly.com

I use rcmd daily (and have just taken the time to dial in the configuration, it works even better now!) and I have to say that it's wonderful. It makes using the computer so much faster and easier: with window scripts too, it's like magic.

There were some teething problems on Golden Gate: keystrokes didn't seem to make it to the rcmd popup, but your recent remediations seem to have solved it. It's a beta OS version too, of course :-)

  • Thank you! Yes, macOS 27 continues to be a pain with the all new rewritten window manager.

    Still working on finding all the edge cases, so sorry if you still encounter problems there. I've been using it since June and still find problems in input handling.

    Like there's this thing, where if an app has Accessibility Permissions and listens to key events (like rcmd does) and then you revoke that permission while the app is running, then your whole system will stop responding to keys and clicks. Until you kill the app in question, but how are you going to do that without a keyboard?

    All apps have this problem, even established ones like BTT, because it's a recently introduced macOS behavior in how the internals of CGEventTap work.

    • Yeah, I thought so -- I imagine quite a lot has changed behind the scenes. It does seem much better than Tahoe, though.

      Yikes, that accessibility bug is bad LOL. It seems a bit like Linux in that way: certain bugs hang around for years. I always found that with Linux, you could either use an LTS and stick with the same bugs for years, or try your luck with a rolling release and trade them for new ones. Anyway...

      I've just started leaning more heavily on the window switcher (moved it to Lcmd) and it's made me so much faster. Had rcmd for ages but never changed it to that; it solves the problem of "where the hell did I put that window?!" when I try to organize stuff into different spaces. It's a better candidate for Lcmd too IMO: I very rarely wonder where I put a random window of an app (and after all, if it's already open I usually want to choose a window), but I'm always switching between certain windows of one app, like Safari. Swiping back and forth between desktops was a nightmare!

I haven't even considered using MCP for locally-running apps! I mean I used to, but they almost always got associated with cloud service access and integrations nowadays, and most local stuff is trivially handled with agent skills.

Thanks for the tips. TIL about BTT having MCP support. Crank looks very promising too, since it looks like a more effective Shortcuts and Hazel replacement

I guess the best way I would put it is that MCP is an RPC for any software in a way that an LLM could interface with easier. Since MCP etc can work with things like Blender.

  • Yep like a self-documenting RPC since you don't have to read docs first to use it. You just ask.

> Get Crank to start Time Machine backups immediately when I connect my HDD and notify me when the backup is done.

Is that not how it works out of the box?

  • Not really. macOS may wait for idle time, may prioritize internal disk and the backup will go extremely slow on an HDD, there's no notification whatsoever.

    It's a very specific thing for me really, I connect the HDD specifically for doing backups as fast as possible then I want to disconnect and store it back so I can keep using my laptop. I don't have a desk anymore where I can keep these things connected all the time.

In the before times, you could use Applescript to expose those things. Is MCP preferable to (LLM generated) Applescript?

This has been what everyone who has supported MCP has been telling people, but coders just endlessly screeched about how CLIs are better.

Everyone on this forum has an absolute paucity of imagination when it comes to applying LLMs to any use case that doesnt involve coding.

  • I think the issue is that any mcp could be a cli and llms are very good and using bash.

    Of course MCP has its use case like if you want auth, or session based actions.

  • Some coders, those that equate being a developer with UNIX, mostly.

    They pay tons of money for hardware, only to use it the same way I was using those DG/UX terminals at the university.

    Naturally there are no coders in other operating systems as well.