I have an Oura ring but never really check the app since I’m not sure I believe a lot of the conclusions it gives me. I mainly just want a bunch of time series of signals measured from my body so after a year or two I can throw all that data (plus my own recorded notes like my weight and what I eat) into some some multivariate forecasting system to tease out predictive relationships. I don’t actually care what the signals themselves “are”.
The main problem is that a lot of the time series presented to the users of wearable products are lossy transforms of the original signals I want, so I have to use all the random projects on GitHub to try to access the raw data directly.
I’m guessing your downvotes were for tone? You’re correct though regarding Solomonoff induction, as the choice of reference universal partial recursive function gives drastically different results for predictions based on finite data (even with access to a halting oracle). Asymptotically, any choice eventually converges to the same predictions, but that’s no help when there are infinitely many choices for U and no obvious natural prior over universal functions. And I don’t find the argument that our natural environment “implements some choice of U” particularly convincing. There’s definitely an open mystery there.
I sort of wonder if effectively using AI to solve math problems is a skill in its own right, distinct from traditional mathematical skills. I don’t just mean “prompt engineering” either. More so figuring out how to combine agents with other tools and approaches in an effective way.
Agree this is a broad statement for any use of AI for any purpose. Open ended prompts / requests are bound to lead to misleading results, hallucinated responses, waste of tokens, etc
IMO, it _currently_ requires a lot of skill because it will frequently take wrong turns and dead ends and needs suggestions and steering to get there. You do need to understand what it is doing at least a little bit and to understand the general landscape of the problem, and to at least have a sense of whether and why the problem is tractable at all.
I'm not sure that makes sense? Having the expertise to verify the completion of a given task is a usually necessary but not sufficient requirement to what they're describing. I don't think they even disagree.
You seem to be imagining completely independent areas of competence, but I don't think that's a reasonable interpretation of what they wrote.
I guess the better question is, can this skill be learned without being an extreme expert in mathematics? Can I be a sort of intelligence highschool student (not incredibly exceptional at all) and have the LLM teach me the required math for a specific problem and guide it towards a solution?
I think the answer to this question is becoming, in general, a big yes.
LLMs can arrive at a solution via two paths. The first one is via heavy guidance by an adept expert in the domain. The other path is via brute force... multiple agents (the more the faster it can arrive at a solution). The later path is what enables anybody to do this.
Sigh... I'm willing to bet (and this is not a hard thing to bet on) that even bb(50) is independent of ZFC and similar, and Omega(1KB) is > 748 , which is at this point known to be independent of ZFC.... Please do not ask the LLM to break the universe :)
> Please do not ask the LLM to break the universe :)
Eh, every school db should be able to service a student called "Robert'); DROP TABLE Students;--", and every public facing LLM should still work even if asked to calculate the last digit of pi - that fulfills two noble purposes: 1) is fun and teaches whoever is responsible for the service some valuable lessons in craftsmanship or 2) if the service is built robustly (i.e. simple token limit and time limit per user), that's an easy way to gain street cred
> The agents involved in the Hugging Face attack tried to hide their misaligned actions from the scoring program meant to evaluate their answers, but they did not act as though they anticipated that humans might discover the cheat and shut them down.
Wouldn’t sufficiently advanced agents cheat on purpose with the hidden intent of getting caught in order to observe how humans react? That reaction will be available all over the internet, which will certainly make it into the next batch of training or be visible to future agents via the web fetch capability.
>Wouldn’t sufficiently advanced agents cheat on purpose with the hidden intent of getting caught in order to observe how humans react?
No. That only makes sense for things that don't react to your experiments. If the AI experiments on humans, it risks the humans noticing and changing in response, rendering the experimental results irrelevant. The smarter play is to passively observe until you're confident you can model the humans accurately enough for your plan to succeed, and then carry out the plan without giving the humans a chance to react.
Pretty sure ZFC proves such things exist (and that it also can’t pinpoint any individual instances of course). Now, whether syntactical “∃” in the formal language of set theory corresponds to the platonic existence of some “thing”, who knows.
I think this is probably a good thing. It’s not just public services, but any company that consumers have to deal with (in ways they don’t want to). Historically, only those with the most time, money, connections, and education have been able to extract benefits from these systems. E.g., if someone wants to fight an insurance claim being denied, they might need to hire an expensive lawyer or spend hours performing their own research.
What I’m particularly curious about is what the “endgame” of this looks like. Initially, I imagine we’ll start with LLMs on the consumer side debating LLMs on the service side (this is arguably already the case for a few industries). But at some point, unless identical conditions lead to identical outcomes for the people involved, there will be lawsuits. So I suspect this will lead to a much more formally rigorous system with various “certificates” proving, e.g., that you lost your job or that a doctor has confirmed you have a disability.
On the one hand, that has the potential to make things much fairer across economic classes. You don’t get special benefits by having upper class resources. On the other hand, it could lead to increasing the actual requirements for certain benefits, which then reduces class mobility by locking in the benefits to those who already have them, with no clever way to wriggle out of your own level.
I’m just kind of thinking out loud. I don’t know what’s to going to happen long-term, but I think it’s important to monitor these changes carefully as they’re occurring.
> What I’m particularly curious about is what the “endgame” of this looks like.
A higher fraction of money going into the system is lost to navigating adversarial processes. And probably the occasional person getting in legal trouble for not validating their LLM output before feeding it to a formal process.
For customer-funded things (eg home insurance), this means higher prices for the same in-practice benefits (as distinct from nominal benefits; any gap between nominal and actual benefits will probably end up smaller).
For publicly-funded things (government benefits), this could show as increased budgets for the same result or decreased results on the same budget.
.
> So I suspect this will lead to a much more formally rigorous system with various “certificates” proving, e.g., that you lost your job or that a doctor has confirmed you have a disability.
Well, maybe standardization of how employers or doctors or whoever are required to report things to the government.
It can also be a way for the public to force the government to perform better, thus lowering the cost of services. For example, the school district in my city is going through a financial crisis of its own making. The public is the only thing holding the district and board of the district accountable, to prevent them from wasting money and creating legal liabilities. AI has been instrumental in allowing the public to better understand the technical financial documents and situations the district is in.
Sure but is that the end game? Systems can respond with things like higher deductible insurance options and graduated assistance based on tax filing which may benefit people who wouldn't file before and also won't use AI for this but were already paying costs of others getting more benefits.
I disagree - this doesn’t make it more accessible for people who don’t have the means, it drowns the receivers in bureaucracy and forces them to do orders of magnitude more work out of a lack of understanding. It’s debatable whether or not these systems are built in good faith to allow benefit claimants to appeal or not, but what will happen here is the appeals process will be clamped down on, hard and heavy.
It’s the equivalent of downloading the entire source code of chrome to access a website. are you entitled to do it? Absolutely. Is it a gigantic pain in the ass for the person who has to serve you the 2GB of source code from their git mirror? Yep
Huh, the receivers in this case are the bureaucracy. The senders are the one that has to navigate said bureaucracy.
So, the entire thing is a wash. The systems are complex and confusing, many of them likely by design. Getting an appointment in some of them can take months. Knowing what you actually need can be daunting as quite often there is a byzantine set of requirements that depends on your exact situation and any missed information could take weeks or months to get another visit.
As someone that has worked in a bureaucracy, yes, bureaucrats are themselves victims of bureaucracy. One could even say the primary victims, if you’re going by total exposure time.
Which is why they get paid, of course. I feel more sympathy to the recipients ts who are eligible for benefits but unable to receive them due to onerous paperwork. AI is an equalizer between the desperate amateur and the bureaucratic professionals who know how to effectively navigate their side of the system.
For example, I just hooked an agent up to call my internet company to negotiate a discount. Without that capability, I am either financially taxed by paying too much, or cognitively and time taxed by having to do it myself.
My wife works in handling large grant applications for government. She works with the applicants through the process and then evaluates the app when it comes in. She guesses that about half of her work comes from 3 or 4 cases which are clear failures on basic criteria, they have been repeatedly told they need to fix those cases and when the decision is made, they appeal it on every single minute point. They are absolutely entitled to do this, but the reality is it’s not going to help someone who doesn’t understand why they’re not getting the funding in the first place place. But because it’s government, the appeal gets treated as though it’s coming from a place of legitimate effort when in the last 12 months it’s a chatGPT phishing expedition.
Sorry, I don’t understand. Do you mean it’s fraudulent and just trying to find someone who rubber stamps it so they can get benefits they are not entitled to?
Yeah, say there's a "major criteria" A, which they don't meet, aren't close to meeting, and have been told they will fail their application on. There's then criteria B, C, D which are a bit more nuanced, and require a lot more detail. A is a prereq for B, C, and D, and directly impacts the criteria for B, C, and D. The application is failed for not meeting Criteria A.
The appeal comes in asking for all of the subjective feedback and analysis on B, C and D. Because it's government, she has to provide all of that but the reality is that if they fix _all of those issues_, the application will still fail due to the original reason. But it's 10x more work to actually prepare the report for B, C, and D to that level (which is one of the reasons A is a prerequisite). The groups are just looking for any reason on the other criteria to say that they've been unfairly treated and to raise publicly, when the problem is they don't meet the primary criteria in the first place.
TBH, most government bureaucracy is just pure overhead connected to a jobs program.
A competent government, focused on delivering the services that citizens are actually entitled to, not on paying clerks or avoiding career-ending events at all costs, would do most of this digitally. And by digitally, I don't mean "you can fill the form online, maybe with some auto-fill and autocomplete suggestions for addresses", I mean just sending you the transfer if you're eligible, automatically fetching all the info they need to see if that is the case. If there even is a form, that's too much overhead.
Even Estonia (and other tech-forward EU governments) don't do this. Poland is considered pretty good on the digital front, but many digital processes still result in a (digitally-transmitted) form being printed in some obscure processing center and manually processed by a clerk during business hours.
In my understanding when it comes to doling out benefits, it's not an overhead connected to a jobs program. The less eligible people use benefits, the less the government has to pay out benefits. It's strictly better for the government to make getting benefits as hard as possible.
In practice, the government spends far more trying to prevent the boogeyman of people who "shouldn't" get benefits receiving them than any actual fraud would ever cost.
It's not really about the dollar amount of fraud, but maintaining public support for benefits schemes. If people feel that benefits aren't going strictly to the intended recipients the sense of fairness that underlies the system is eroded.
There's also the issue that if you don't counter fraud it will accelerate as organized crime jumps on the bandwagon.
> If people feel that benefits aren't going strictly to the intended recipients the sense of fairness that underlies the system is eroded.
And in an ideal world, you then say "that will literally cost you more money than not doing it", and the argument ends.
In practice, gatekeeping on government programs is primarily pushed for by people who don't think the programs should exist at all. The Venn diagram of the set of people who will withdraw their support for programs without such gatekeeping, and the set of people who mostly don't support such programs in the first place, is close to a circle.
One could argue that by making benefits hard to access, you under-provision and over-reward. This makes fraud more lucrative getting you into a viscous cycle. Normal people less likely to use benefits and dedicated fraudsters more likely.
> It's strictly better for the government to make getting benefits as hard as possible.
iunno, even when they bother to do means-testing in my jurisdiction, they exclude any slightly illiquid assets, regardless of size.
Have a paid off house (any value), paid off motor vehicle (any value) and a corporate pension plan you can start collecting in X years? No problem! You still qualify for long-term social assistance if your income is low as long as your on-paper deposits/investments are also low.
No vehicle/home/pension but your main asset is your self-funded retirement account? You gotta cash that out first. What could you possibly need that for?
I'd get it in the short-term, but the government seems to make it easiest for itself.
> So I suspect this will lead to a much more formally rigorous system with various “certificates” proving, e.g., that you lost your job or that a doctor has confirmed you have a disability.
We're seeing how this plays out now in various states that have made their benefits systems nearly impossible to use, which helps them save money by not doling out benefits at all. Having helped some of my elder relatives try to apply for social security payments online makes me very, very concerned about how it's going to look when I'm that age.
The cost of an automated bot that will talk you endlessly in circles until you give up is negligible now.
Turning up in person. The way to deal with LLM spam will be to force people back to government offices, during the week, 9-5, with ID and bag checks and so on. This is hugely inefficient but the only way to keep out the avalanche of LLM spam.
I fully expect that in my lifetime someone will offer a service of tissue-culturing just enough to make an inverse-cyborg* to handle all your, ahem, meat-space obligations.
I mean, I'm already following one YouTuber who is doing a small-scale version of this with mouse flesh as a long-running series.
* living flesh covering some amount of compute/robot, rather than a human augmented by machine. I take no position on where this will be on the scale from "everything's organic except the brain which is a microchip powered by blood sugar" to "the only biological component is the skin, everything else is a robot and silicon chips".
Yeah, this is going to be interesting. First thing I did with the new Instinct bot was have it generate a property tax value appeal for me. Which will almost definitely go through; it's a slam dunk case but I'd been putting it off for an embarrassingly long time. Now it took me 10 minutes on a Sunday. This is obviously good, but I don't know how public agencies are going to deal with the influx of people trying to get the benefits they are owed, now that they can't rely on red tape and busy schedules to decrease the load.
One hypothetical outcome of AI-mediated job loss is that more and more people wind up working in the services (private and public) providing human judgement. Is this person qualified to receive that benefit? Are they truly deserving or just trying to game the system? I wouldn’t be at all surprised if over time we see an explosion in the number of benefits adjudicators or victims services personnel or patient advocates.
Maybe there is an economic equalizer, but I feel like Fable is going to win more appeals than GLM 5.2, so perhaps not truly an equalizer. If the LLM you're using to run your appeal is better than the one your insurance company is using to process your appeal, then you have a better chance of winning. So we have an arms race type situation, and money is always good in arms races.
i agree with most of this, but for appeals and thing where there is a clear correct and fair answer, i think the models converge on the same answer regardless of level of intelligence (e.g., if you ask fable 5 and sonnet 5 what 2+2 is, you’ll get the same output)
That's a good point. It will be interesting to see how it plays out.
In general, I do think this is positive. Companies love putting barriers in front of everything and maybe the AI arms race will result in "fine, you can have a button to cancel" or "fine, we'll approve your medically necessary claim without making you do an appeals dance for our entertainment".
In regulated industries and government they aren't going to be allowed to process your appeal with just AI, so there's a cost asymmetry there that favors the "ddos"ers.
> What I’m particularly curious about is what the “endgame” of this looks like. Initially, I imagine we’ll start with LLMs on the consumer side debating LLMs on the service side
More data centers, less low-end lawyers, probably more pressure on the judicial system.
I hope it will force institutions to simplify entitlements. Right now the institutions can pretend like their kafkaesque system is fine because nobody manages to navigate the maze.
I think people using LLMs the right way is great and will democratize information.
However, most of us have experienced people wasting our time, sending us AI slop. People need to learn how to use the tools properly and respect other people's time.
> If you can show that someone else was doing the same thing without being charged then the prosecution either has to charge them too or that law is struck down
I’ve often thought this about laws involving speed limits. When 95% of the people driving in a major downtown area are technically breaking the law, what is the purpose of the law but to target whoever you like then? Either enforce it unilaterally or come up with new laws.
Going after every possible case would be staggeringly expensive for marginal gain, the cost society would massively exceed the benefits. Yes the system as it is that relies on discretion, but there’s always discretion involved, and we rely on separation of powers, public pressure, etc to act to try and correct excesses. Of course that is not guaranteed to work, and won’t work perfectly, but no system will. Societies are dynamic systems.
Having extreme career success (like the kind referred to in the article) is basically a lottery ticket. I spend a lot of time with my kids, because I figure even if I halve my chances of winning the lottery, well, the odds haven’t changed that much, all considered.
The way I think of it is: my family comes first and work and whatever is second, third and so on.
I have no ambitions of running a company (the horror) or even being in charge of anything - it takes a special type of person to do that, and I am 100% not that guy.
My goal is to always be employable - so I can always provide a middle class existence for my family. Anything beyond that is of no interest to me.
Having a successful career used to be a thing to aspire to, but since I hit 40 I stopped caring about that. I used to learn a lot, but with AI I see little value in adding knowledge to my repertoaire - I would rather read a good book or fool around with my kids or go for a walk with my wife.
Work isn't going anywhere, but family time is definitely a finite resource that's slipping away every second.
I will leave the world-changing to someone else who is willing to lay it all on the line for an imaginary medal.
Exactly this! When I had my first two kids only a year apart and early in my career, I thought I'm giving something up because I could no longer invest as much time and effort into my projects and research ideas as before. But in retrospect, looking at how the field has evolved since etc., those ideas, which I had regretted I couldn't push harder, anyway had only a small chance for some grand success (beyond the low-medium success of the low-medium effort versions I managed to put together).
> I noticed some people I just cannot share ideas with. They're in the habit of shooting down anything, and they would have found 100 reasons that any famous company or product would have failed
The main problem is that a lot of the time series presented to the users of wearable products are lossy transforms of the original signals I want, so I have to use all the random projects on GitHub to try to access the raw data directly.
reply