Comment by 1vuio0pswjnm7

1 year ago

.

   #!/usr/bin/env bash
   #
   # memo(1), memoizes the output of your command-line, so you can do:
   #
   #  $ memo <some long running command> | ...
   #
   # Instead of
   #
   #  $ <some long running command> > tmpfile
   #  $ cat tmpfile | ...
   #  $ rm tmpfile
   
   to save output, sed can be used in the pipeline instead of tee
   for example,
   
   x=$(mktemp -u);
   test -p $x||mkfifo $x;
   zstd -19 < $x > tmpfile.zst &
   <long running command>|sed w$x|<rest of pipeline>;
   
   # You can even use it in the middle of a pipe if you know that the input is not
   # extremely long. Just supply the -s switch:
   #
   #  $ cat sitelist | memo -s parallel curl | grep "server:"
   
   grep can be replaced with sed and search results sent to stderr
   
   < sitelist curl ...|sed '/server:/w/dev/stderr'|zstd -19 >tmpfile.zst;
   
   or send search results to stderr and to some other file
   sed can save output to multiple files at a time
   
   < sitelist curl ...|sed -e '/server:/w/dev/stderr' -e "/server:/wresults.txt"|zstd -19 >tmpfile.zst;

Those commands are a (1) harder to grok and (2) do not actually use the memoized result (tmpfile.zst) to speed up a subsequent run.

Can you give a more complete example of how you would use this to speed up developing a pipeline?

If provide sample showing (a) input format of text and (b) desired output format of text, then perhaps can provide an example of how to do the text processing