Comment by open592

1 day ago

Let's hypothetically say I'm a PHD student who is half way through my studies and I have a halfway written version of one of these "preprints" - what do I do?

Seems like a lot of PHD students are doing to have to pivot the entire structure of their PHD studies? Or just produce something which is already written by OpenAI?

As others said, it's not that this isn't a phenomenon that is unique to now. It happens. I had to pivot a bit of my dissertation near the end because at a random conference I spoke with a researcher from another continent and realized one of my ideas was already out there in some form. I just missed it since it was in a conference proceedings outside the usual set I looked at. So, I had to scramble to adjust and still come up with something novel. I survived, and defended, but it wasn't that much fun at the time.

What isn't so normal is the probability and ease by which this kind of thing can happen today versus decades ago when I was in school. As OpenAI said, it only takes a few hours of compute to do what likely was much more than a few hours of human effort. The only reason this kind of scooping/overlapping was rare was mostly a function of how fast other humans could do the same work. With machines, that totally changes the relative pacing between the human trying to learn how to be a researcher and the machine that can grind out results.

I'm less worried about the phenomenon of overlap and scooping and such. I'm more worried about the long-term impact on fields (not just math), especially considering the early stage students and researchers entering the pipeline now. I'm not sure what happens to disciplines when that pipeline stalls.

Let me give you some perspective: my entire phd was made obsolete by CRISPR. It was a wonderful thing.

  • Interesting. If you don't mind, could you please share a little bit about that ? You already finished your dissertation ? It was about the works of Doudna and Charpentier ?

    • No, back in the late 90s and early 00s, people were trying to engineer custom nucleases and transcription factors, my work was on doing molecular dynamics simulations to optimize TF sequence specificity (similar to engineered zinc fingers) for gene therapy. I wrote up my dissertation and published it in 2001, and then went off to find enough compute, IO, and smart people to make it happen (https://research.google/blog/groundbreaking-simulations-by-g...).

      My approach would require custom engineering for every different sequence we'd want to target. With CRISPR, you just "program" the system with a guide sequence, you don't need to do massive engineering to solve a protein design problem.

      3 replies →

  • For every door that shuts another one opens - great if you're not a cabinet-maker.

I have had this conversation with my PhD students yesterday. I am 100% sure that all of their problems can be solved by publicly-available models now (I solved a case of one myself as a test, it took 15 minutes). So the challenge for them is to see how much they can accomplish in their allotted period, and still pass a defence on at the end of it all. The PhD defence is going to become all about a test of understanding, not a test of quantity of publication.

A much bigger issue is: What is the point of any research mathematician publishing anything now? I really hope that one positive effect of all this will be to finally topple the awful peer review model we currently have, with the biggest publishers gatekeeping with extortionate fees.

This happens all the time, even without AI. Other researchers or PhD students can publish the same results before you. I say that based on my experience during my PhD.

  • Yes and no. It's pretty normal for students to get scooped, sometimes even multiple times (ask me how I know...).

    It's much less common for students to have their entire thesis direction removed from under them, as might be the case for someone working in fine-grained complexity assuming 3SUM has no subquadratic algorithms, or working assuming ~UGC. Both of which (publicly) seemed like perfectly valid research directions until a couple hours ago.

    There might be stuff to salvage from their conditional results anyway, but this is not your average scooping.

Use whatever is published as the new base for your research. Use AI tools to help you going forward.

  • In other words, you've already taken a gambit with the first half of your PhD, now take a second gambit, praying that you have something to publish by the end of your PhD.

    It has to feel awful to be in this position.

This has always been a challenge for PhD students and researchers, it's just far more likely to occur now it seems. Getting scooped doesn't feel good, but it's a signal you've been thinking about things other people care about.

I think eventually companies won't get as much stuff that's usable for marketing, so they'll stop investing so much into ai for math, so eventually cheap and poor graduate students will be able to do relevant work again without worrying about getting scooped by a company with a million GPUs :-/

Perhaps consider switching to a more applied field. Experiments in the physical realm hold value, especially if you document the process.

The purpose of a Phd is to train researchers, not to solve one small problem.

Some math PhDs would spend a year or so doing things with AI and lean, and graduate. And keep doing more math afterwards.

Some others, with more stubborn advisors, will keep trying to find a gap where there's no AI progress.

CS subfields go through this every ten or so years.

If you had halfway to one of these papers you would be anyhow be in the top echelons of math phds so probably you have less to worry than most!

The follow up question then naturally would be how do phd advisors with people whose fields are in someway premised on making breakthroughs in theoretical fields that AI can solve work deal with it?

At my (German) university, the solution would have been to change from a "cumulative" thesis (which requires peer-reviewed papers) to a monograph-style one. Because the general rule was that your undertaking must be novel ar the time you submit your thesis topic, not necessarily at the tome the thesis itself is submitted (years later).

> what do I do?

Very hard question.

Your work makes you one of the very few people who really understands the problem and solution and its significance.

It could still be interesting if your approach to the problem was different to theirs.

Don't worry, if you use openAI and get lucky they might offer to share credit with you for your work.

Don’t paste your research into these AI because they will train on it and then scoop you.

  • I think that happened after word of the project reached OpenAI and they allocated resources to it.

> what do I do?

Precisely what all NLP researchers and the ML community at large did in the last few years: embrace the frontier and realize that attention is all you need.

Ask OpenAI for money when you don't have a job or future?

I guess the only answer is to adapt with the tools. If we can't do that, then yeah, we're in trouble.

Frankly, I believe PhDs don't have to be novel in the absolute sense, only new research to the student. I can't remember how many times I've effectively invented something from a clean room approach only to realize someone already published something years ago or that my "new" algorithm has a name. So this really shouldn't be holding anyone back.

That would be unfortunate but the world does not owe you anything and is not going to stop for you. Which is also a valuable thing to learn in your 20s (ideally earlier).

> Let's hypothetically say I'm a PHD student who is half way through my studies and I have a halfway written version of one of these "preprints" - what do I do?

I did get my PhD...before AI. And my honest advice (to myself back then, even) would be: quit the PhD, become an electrician, and work hard to buy a tiny house in the middle of nowhere to watch the world burn in this madness.