Sunday, April 26, 2020

The legend of the book of the film of the record of the poem of the graffiti of the urban myth

"Which came first, the egg or the chicken?" is the perennially annoying question that consultant philosophers use to impress naive clients.

A much more serious question is why people hate the film  of the book, or  love the film, but  hate the book  it came from, or love the film, and hate the book  that came from it.

To solve this problem, I think we need to carry out a large scale analysis based in causal inference  (why not) .  Clearly we have the equivalent of the adaptive clinical trial or the series of unfortunate natural experiments to choose between. We can start with some obvious candidates.


  1. The Godfather
  2. Paddington
  3. The Princess Bride
  4. The Hobbit
  5. 2001 or The Sentinal
For each of these we need to look at all the possible features that could lead  one to prefer a book or a film (length, sentence structure, plot, character, year published, alphabetic position of author directors name, jokes, box office takings, time in best seller charts, influences, sequels, spinoffs, live action, cartoon, comic book, illustrated, etc etc) and build a bayesian model of how one moves from one state of mind (Hobbits are boring) to another (Smaug is cool), or from one opinion (why did no-one ever actually read the princess bride) to another (You keep using that word, etc etc ).

Two dominant theories to date are

  • The Ordering Theory
  • The Gap Theory


Then we will finally know the truth.

Wednesday, April 22, 2020

cat > /dev/kb


I'd like to bring up a very important topic for all my inky-fingered friends (and I am not
referring to my experiments with spilling Quink on the ebony fretboard to see if I can play faster).

The cats. come and sit. on the keyboard. in front of the screen. while you're trying to work.

How can we fix this? 

Well, here's my patented invention that I think is going to work.

You know how the keyboard layout was invented by Benjamin Franklin for President Theodore Roosevelt when he was digging the Panama Canal? It is an oft told tale - the problem was that all the typists producing  reports to send back to Washington DC had to type an inordinate number of lower case 'a's every time just to get the headed paper to look right. For this reason, they moved the keys for 'l', 'n', 'm' and 'p' as far away from the 'a' as they could get them. Back in those days, most people were one or two fingered typists so this slowed them down, so that the 'a' key stopped breaking, leading to incredible misunderstandings between the  rmy, nvy and rfrce ('o' was a lesser,but not insignificnt prblm).

So the new keyboards worked a treat until they started working with a French company who still used the old Aztec name Pznzmz Cxnxl, which led to the second new keyboard, which nous aimions tres biens ces jours. and so on and so forth.

As code breakers amongst you will remember, (or possible Sherlock Holmes stick insect fans) the problem is to do with the popularity of different letters f the alphabet in different tongues. Why the alphabet is in alphabetical order is an interesting question which I'll leave for later  (just noting for now that it isn't, for example, in Arabic, Hebrew, Greek and Cyril Smith's languages, the third letter is 'g', not 'c' - go figure how the Romans got that wrong along with really poor ways of counting). And of course, Scrabble scores - which are in the opposite order to the popularity of letters ( a bit like Nathaniel Hawthorne's novels).

So "how does this solve a problem like a cat?", I hear you ask, Maria.

Easy peasy. we place bigger springs under the unpopular letters and so when the cat sits on the nice warm laptop keyboard, instead of getting a gentle purr like vive from the fan, it gets prodded uncomfortable in random places.

This will also let us revert the keyboard layout to being alphabetic, since it will just be harder to depress the unpopular letters which will slow us down between popular ones.

Of course we could confuse the cat further (as if such a thing were possible or even desirable) by choosing springs from a French or dare I  even suggest, a Chinese (Mandarin, not Catonese (sic), of course) layout of spring constants. By Hooke or by Crooke, we will have to solve the problem that the Chinese layout would require at least 5 springs of different strengths to operate really successfully, but we believe that yhe market for this in china will be as big as that for Dragon Nets.

I will be inviting investors to my alpha-beta-gamma product launches shortly, meanwhile I leave you with the experimental result that  you may wish to try as well, in these distracted days. You can teach your cat a foreign language easily. I have ours completely versed in French - when I say va't en or viens ici, she behaves in exacfly the same way as when I say "get the 'f off that keyboard" or "where are you didums". This does not work with dogs. In fact I know several dogs in the Dordogne that response to "Get off my leg fido" in exactly the same way as if you offer them a biscuit.

Science is a marvellous thing when used carefully.Electromagnets more so - in the new version (delta-key) of the boards above, we replace the springs by electromagnets and now can use this to train humans to type faster, and untrain cats to sit on the keys. Gnu Emacs key bindings will be available shortly.

References

Keyboards:
qwerty or azerty

Frequentists:
code breakers and scrabble

Dvorak:
chording

Medication:
While waiting for the coffee to brew,

Monday, April 20, 2020

ethics, policy, regulation and contact tracing apps - babies and bathwater

0. How to get it right: Harvard ethics review/roadmap to pandemic resilience

1. if you are going to criticise the ethics of contact tracing app work, first establish a baseline - find out how manual contact tracing is done, what data is kept, what triggers it, what privacy risks there are. Who does the work? are they trained? is there a log (to avoid duplicating contact notifications by multiple staff and to record the test status of people). How does consent work from the pattient who's just tested positive and is probably stressed? How many false positive and negatives are there (people they falsely remembered they'd met, or encounters they forgot) etc etc- this is the standard an app has to meet at least.
2. know your medical ethics - it is standard that a new technology is introduced provided it is at least as good as existing "treatments" and no worse - see above. If contact apps tracing is also faster, and therefore reduces the number of people the virus spreads to before all possible infected people are found and isolated, then factor this in, as it is part of the care requirements.
3. Don't talk about stuff you don't understand - i've seen vague criticisms of the use of BLE (Bluetooth Low Energy) as it isn't "accurate enough" without a single citation on measurements that support this. There are multiple measurements and techniques that support that it is an ok proxy for encounters between people carrying capable smart phones with the relevant app running. The main criticisms are i) that might only be 60-75% of phones (depending on country/region/demographic) and ii) only around 75-80% of people have any sort of smart phone. See 1/ its additional, and faster/complimentary, not a replacement for manual contact tracing. Also see the care taken by Google/Apple (see prev blog) in terms of taking care of privacy- this is largely better at protecting people than manual tracing can be (there are modest exceptions - exercise for reader, think of one).
4. No-one's claimed you use contact tracing to replace testing (or more ludicrously, to replace the hunt for treatments, or vaccines or the actual provision of PPE for key workers who encounter a lot of potentially infectious people (including care homes, bus drivers, supermarket checkout staff etc). Don't claim people want to re-prioritise resources because they are techno-solutionists without actually finding out their motives and community. The idea of contact tracing apps came from (and is supported by) epidemiologists (e.g. from LSHTM in the UK). Tech people worked from what they asked, not from some bluesky fantasy. This includes the empirical testing of which is preferred proxy for encounters, and the fact that it is predicated on testing. Testing alone doesn't fix things fast enough either (unless you had a 15 minute test cheap enough to run on everyone nearly daily).
5. contact tracing doesn't have to be in more than a few percent of the population to be useful for its original purpose, which is to get more precision about the epidemic parameters to improve models, learn about asymptomatic carriers, infection rates between groups like children-to-adults, and the expiry date on immunity, whether acquired through surviving infection or from eventual vaccine deployment (many vaccines also have limited lifetime though usually better and longer than having had the disease, we still need to know).
6. Mission creep:-  discussed elsewhere - a mix of tech, regulatory and legal frameworks need to be clarified to minimise this risk- including (obviously) sunsetting.  This is not new.  People that work in clinical trials/medical ethics know this stuff - if you are a tech ethicist and you have not read up standard protocols in that space yet, please do so before criticising tech app writers who have. The goal of the privacy preserving/decentralised bluetooth API from Google/Apple is not to mess up earlier more centralised app designs, it is to offer a more ethical way forward and represents the way the tech sector has considered best practice ahead of some of the people criticising them

What's worse than techies who ignore ethics and context? ethicists who ignore the tech and context:

For the avoidance of doubt, Do No Harm.

Comprehensive list of tracer apps, initiatives, design docs etc

Friday, April 10, 2020

Some DP-3T & Apple/Google contact tracer abuse questions...

Contact tracing plus testing is a hope for getting out of lockdown, once we are well past the current peaks in the Covid-19 pandemic . Lots of apps have been proposed, some shipped. Most recently, privacy preserving apps have been designed in response to fears about misuse of the contact data. Apple&Google have specified an open API&Service for bluetooth low energy contact tracing with privacy. It looks like a good fit, technically, to some of the newer app designs. It does (a little) remind me of what adding privacy to WiFi AP scanning did (to prevent revelation of all the places someone had been by eavesdropping the list of prospective APs in their scan), but to a very different end and in a different way - see links to specifications below. Some comments added on NHS proposed app at the end now.

People are concerned about how this might lead to privacy invasive apps in the future, but first, why do we want this now:

Aside, to keep an epidemic in "virtual lockdown" you need to able to trace and isolate cases before they infect further people and restart the epidemic exponential growth ahead of your trace rate capability. This means there's a relationship between the reproduction rate (R0) of the epidemic in normal population behaviour (contacts that might lead to infection) and the fraction of people likely to be able to give fast accurate contact information - with nominal R0 around 2, this is estimated in the range 40%+ of people out and about. If people wear masks and observe social distancing, the baseline R0 might be somewhat lower. With proactive testing (random or periodic) you also trigger things earlier for people testing positive so the effective R0 is then even lower - the goal is to keep it always effectively well below 1. But note the number  40% of UK population (or even just households) is 20M (10M) roughly.

Could you build an app to "round up all the co-conspirators"?
or all people that were at this protest at this time with this person?

1. agency (replace healthcare with bad cops:) coerce person to equivalent of test +ve: sends notifications: 
2. agency coerce people to reveal whether notified or not 

Could latter be required by, say, employers (e.g. good ones like healthcare, or bad ones like xxx)?
How is that new compared to current Real World contact trace/notify done through interviews/phone visit

Firstly, service doesn't give precision time, nor is their geo-location as part of it.

Phones may already potentially separately run geo-location, so not clear this adds a lot apart from additional evidence of co-location, and spatial precision. So if any of the phones in a co-lo event are also reporting position, you "infect" contacts with a possible inference, if someone can coerce ALL possible contacts to reveal presence or lack of notifications...obviously people at protests could turn off service, and later not ask for notifications. Would that then be evidence too? This seems like a pretty complicated and far fetched scenario...Not very good evidence that some people out of 20M might be co-conspirators. Not clear how the coercion scales without becoming somewhat visible.

Explainer/proper use case/reference:
https://ncase.me/contact-tracing/
Google/Apple BLE explainer
https://www.blog.google/documents/57/Overview_of_COVID-19_Contact_Tracing_Using_BLE.pdf
Tech spec:
https://www.blog.google/documents/56/Contact_Tracing_-_Cryptography_Specification.pdf


Pandemic mission creep "best intention" temptations:-

1. "Self-report" and Test certification verification.

Given the trigger for the upload of crypted contact info is a positive test with authorisation by the health authority, there's a strong temptation to bundle test certificates 
+ve/-ve/timestamp/ virus v antibody, into an app...

This is orthogonal to the contact side. but employers (especially healthcare employers) might require verifiable clear tests for staff (like CRBs for teachers etc). Is
failure to do something about being notified also a breach of some employment agreement? Is commerce going to coerce?

I suspect people who work in jobs that you care will actually want to respond to notifications and  get tested too, so can tell to self isolate/get treated/get better and back to work in safe knowledge, so
incentives are aligned, no?

In the NHS app case, there are two separate triggers for using contact history to send notifications: 1/ is a self report (yellow alert), 2/ is a positive test result (red alert). A colleague suggests that there should be an intermediate trigger where a call to the UK's 111 service that results in suggestion to self-isolate, could be accompanied (like the positive test result) with an authorization code to the app (given over the phone to the subject) so that like the test, this trigger (say amber alert) would be much harder to troll with fake self-diagnoses and might act as a deterrent to such behaviour since the 111 caller would be identifiable. re-linking with the patient is no more risk than it was in the test case, either.

2. Isolation/lockdown location compliance

Since we don't have absolute geoloc at all, is there a way to find if notified people were in contact with a person who was infected and in breach of isolation/lockdown rules, more than current Real World contact tracing would reveal...? This seems not to be made easier by these 
contact tracer approaches. See above. 

Other concerns include false positive rates in self-reporting - this applies whether the data is centralised (NHSX current app design as of 12.4.2020) or decentralised as with the Google/Apple/DP-3T.

We can assume that there will be fairly high levels of people stressed in the current lockdown, and potentially experiencing some symptoms (e.g. coughing at the slighted thing). We're currently heading out of the period of seasonal flu, so people having genuine symptoms, but caused by something less risky, will be in smaller numbers perhaps? Nevertheless, this is going to contribute a significant "false positive" rate. However, given the goal of all this tech (coupled with more wide scale testing) is to be able to leave lockdown, the effect would be to have some larger number of people self-isolating than expected, but a much much smaller number than the current 65M people stuck indoors. It remains to be seen what that rate would be, but even if 5 times the rate of real symptoms, this would (after the current peak is over - say early May) be quite a modest number. And it is "failsafe"

Another threat sometimes claimed to these systems is trolling. This I don't buy. The whole point of the bluetooth scanning algorithm (since we did ours 11 years ago in Fluphone) is that someone would have to stand next to you (less than 2 meters away) for 15 minutes (or so) to trigger adding you as a contact. You'd probably notice people doing that in the supermarket, on the pavement, etc. Fleeting encounters are not triggers. That's part of the design.

The third criticism I've seen of these contact tracer apps is that they need a significant fraction of the population to run them for them to "work" - actually, this is not strictly true - they need a significant fraction of an infected person's social group (friends, family, colleagues) to run the app to help. This is true for the contact tracing side. but all contact tracing is partial - it is an attempt to reduce the reproduction rate of the epidemic below 1 - any contribution to that reduction helps us avoid a second wave.
The app is also useful (as discussed above) for gathering details to build a more precise model of the epidemic, mathematically, so things like pre- and asymptomatic carrier infection is characterised, and the rate of child-to-adult is understood better, and even the expiry of immunity. For that to work, any reasonable number of people running the app will help. Given other apps (eg. Zoe/Kings app, or the Covid-Sound app have seen  thousand of downloads a day, it is clear that reaching a decent target for that purpose is achievable, whereas to get to herd-levels of contact tracing coverage many be harder.



Question: apple are saying they will mandate the use of the new privacy/decentralised bluetooth scanning API for IOS devices to run scanning in background - is this already in place, or is it after they (and google) release the new scanning code? Would a centralised-store app like the NHS one be blocked (either from release through the Apple App store, or further, actually unable to run (in background) on IoS devices right now? Will update this soon as someone upates me:-)

Baseline: how does manually tracing contacts (extracting addresses/phone numbers from a tested person, and subject to imperfect recall, possibly including people they didn't actually see and forgetting ones they did)- how is that better than digital contact tracing in safety&security?


meanwhile, also, some data on manual versus app based tracing impact on reducing R0 - from lancet paper based on data from china

Tuesday, April 07, 2020

covid-19 & interventions - very very speculative

looking at Mark Handley's graphs of many countries evolution of the pandemic, and the interventions, i'm going to engage in some idle speculation - please don't take this seriously or as a prediction - its just thinking out loud...


people see family multiple times a day
people make daily trips home->work/school
people make weekly trips (home->shops or trips to country)
people make monthly trips (business -> other countries)

this pattern of self-similar journeys underlies this study of cinter-ontact intervals and duration.

If you look at clustering starting in Wuhan, and then to rest of china, then to other country, it really looks like that. Rhyhm and randomness in human movements as also explored in this ref paper.

It seems that in a trip you might meet 1-100 people but only infect 1-2 so its quite hard for most people to catch, but somewhat easy for some
(severity is a whole other question, maybe, but maybe its related too)

Baseline:
35% daily increase corresponds to a doubling time of 2.5 days
assume basic R0 is 2-3 -i.e. as above, so each person adds 1-2 people a day
but they aren't necessarily infectious for 3-5 days so you get double every 2-3 days

Social Distancing
22% daily increase corresponds to a doubling time of 3.5 days
distancing works weakly - i.e. your infectious person travels
more carefully in their daily trips but not carefully enough. see this LSHTM paper for more info which looks consistent?

Speculate - if Covid-19 can be spread by touch, this might be
evidence that 2 meter is fine if everyone washes every time they were near where an infected person touched. so maybe combination of masks and washing would be as good as lockdown if 100% observant, but if 50% obvservant, only roughly halves the spreading rate.

Lockdown...works:
13.5% daily increase corresponds to a doubling time of 5.5 days
lockdown week 1 < works, but you've only removed half the people in the first week

8% daily increase corresponds to a doubling time of 9 days
lockdown week 2, remove other half....but still have tail of who was infected 2 weeks back, showing symptoms this week

0%
lockdown week 3?

Saturday, March 21, 2020

epidemics and human contact statistics...

way back when, my colleague Eiko Yoneki built the Fluphone project, which has suddenly become rather relevant again after a decade. Some folks in Singapore have nicely done something similar with good privacy properties! NHSx are on it in the UK, as are lots of other people...

we had earlier studied human contacts - here's a paper by Augustin Chaintreau et al based on work in the Haggle project which involved empirical studied of how the time between and duration of encounters between people is distributed.  Some of the data sets are freely available in the excellent Crawdad repository.

When modeling an epidemic (e.g. to figure out whether it will collapse, sustain, or go pandemic), people start with the SIR model (susceptibility, infectiousness and recovery) - this basically leads to classic S curves over time for the number of people infected (or fraction of the population) - the other important number quotes is R0 - the number of people each infected person passes the disease on to. Recovered people usually have some level of immunity, so they are "removed" from the population, and not typically returned to the pool of susceptible people (at least for some time, depending on disease and person).

Some problems with naive models:
the values for S,I,R are population averages. As is R0. In fact, R0 obviously varies over time, as the number of susceptible people an infected person can meet must typically decrease.

In fact susceptibility for a disease also can be influenced by prior incurred immunity. This can vary with age, gender and other factors, inherently, or because of similar prior infections, or because of vaccination.

As can infectiousness (simple example - if you are asymptomatic carrier of a disease spread only by coughing, then if you don't cough, it is hard to spread it.).

The point of the encounter data above is that that also isn't simple. People have varying levels of popularity ("degree") and centrality (they are on the path between more or less friends (of friends (of friends (of friends)))) to the sixth degree and more. One study of this by Watts et al shows how this leads to multi-scale, resurgent outbreaks. Ground truth needs us to test everyone!

Vaccination programs often target the vulnerable groups (seasonal flu), but sometimes target the whole population, but possibly indirectly - e.g. if you vaccinate enough kids against measles, mumps, chickenpox, polio, smallpox, and it lasts til later life, eventually there's almost no-one in the population left to spread the thing anymore. The vaccine can be made from a weakened version of the disease, which then "teaches" the human immune system in a way that lets it respond appropriately to the full-strength one later. Herd immunity normally refers to having enough of the population vaccinated that the SIR model tells you the disease won't get anywhere far before only meeting immune people.

The contact interval/duration models are power law and clustered. This tells us that they are driven by heavy tailed distributions of popularity (aka rich get richer, in trading terms), but also affinity (e.g. kinship, friendship, work relationships, etc).  For social distancing to work, you need to combine breaking all these kinds of links, social, work, entertainment, even (actually, especially family if you have a tight knit generational spread!). but also you need to target high popularity or high centrality people especially - this has been used by people in a wide range of areas such as offering advice on safe sex to sex workers to reduce HIV spreading, and even smoking and obesity (e.g. see Christakis work there).

The point of this blog (which may contain errors-  please send me corrections if you spot any!) is to try to explain that it is relatively complex to deal with real outbreaks. When we have phase changes in epidemics, and power law coefficients, some small changes can really fix things, other changes have to be very significant to make a difference. We need adaptive, and (as seen here on heterogeneous) responses. Above all, we need continuous, and continuously more accurate, measurement of all of the above to do this adaptation precisely.

Saturday, January 25, 2020

An Architecture for Spread Spectrum Computation




This is an unsuccessful proposal to Facebook about an
an intermediate Instruction Set Architecture for Spread Spectrum Computation. We target nano-services constructed 
from lambdas as a backend from an intermediate system, to allow for fine grain, and elastic, fault tolerant 
computations. Was an extension of an earlier idea by Steve Hand.

We believed it fitted in their research call topics on
Scalable, elastic, reliable distributed;
Programming languages&compilers for platform agnostic; and
Resource provisioning for efficient ML.

I guess it was slightly too ambitious:-)

Saturday, January 11, 2020

some memories of Peter Kirstein

Peter Kirstein, who passed away earlier this week (8.1.2020) was my PhD advisor, back in the 1980s. I joined his research group fresh out of the MSc there in 1981, and was working with Rob Cole who ran the collaboration with RSRE (Royal Signal and Radar Establishment, Malvern) who connected via the UCL gateway on to the Internet.

Peter had been at this game for quite a while (records say 1973 was main ARPANET link, with the connection to NDRE in Oslo). Peter had gathered a team of people both for his own research program, and to deliver undergraduate and Masters courses in Computer Science, having recently founded the department (actually it was part of UCL statistics in the Pearson (yes, that Pearson) building, and then separated, as CS grew. But at least til 2000, it remained in the very nice location right at the entrance to the main UCL quad.

Peter's style of management both for the research and for the department was very collegiate, and entailed quite a bit of delegation - to that end, for admin and teaching he'd gotten some very competent people around him who did a great job -- this instilled responsibility in people.

When I joined, Peter not only had the main research project in Internet related work outside the US he had great links on into Europe, wth collaborators at CNUCE (CNR in Pisa now), NTA in Oslo, FGAN in Germany and others, and also not just the defense link with RSRE in the UK, but also a big interconnection project with Cambridge, Loughborough, and other universities as well as Logica using Cambridge Rings (10Mbps before ethernet - UTP, not coax:-) with wide area satellite links, both for the US (SATNET - the Atlantic Packet Satellite ) and the UK (I think Stella?).

What this tells you is that Peter thought big - very, very big. And this bought challenges in terms of technology, policy and management. Folks in the US programmes (e.g. at UC Berkeley, and LBL)  found this interesting - for example, the satellite link had a very high latency (.72 seconds) compared with land lines (terrestrial point-to-point cables from telecom companies "the phone" trunks). The link also had unusual errors/losses - i remember seeing really bad performance one day and puzzled, asked a colleague who pointed out the window at a thunderstorm/lightening...but also to get funding for this scale of work was a coordination problem with multiple agencies ("stakeholders" is the trendy term now) from US, Canada, UK, Europe, government, industry, academia. Policies collided - another challenge - how to share a network between different agencies with different funding and collaboration rules? Policy routing (BGP) emerged - folks at MIT were instrumental in extracting the policy rules to see how one might build an inter-domain internet.

Peter was on top of all this, thinking about how to drive forward to the next problem, and showing incredible patience with some of the partners who took years in some case to understand what was needed.

One of the things  helped Peter with was the marvelously vaguely named International Collaboration Board, who actually had a charter for a while, which just said "the purpose of the ICB is to hold meetings".  It was actually the vehicle for the resolution of some of the challenges. We also did a fine line in drawing network maps, sometimes down to specific hardware details of line cards (e..g with BBN folks) and other times just scribbling the now ubiquitous "cloud" image (i.e. abstracting away all the (un)necessary details...).

Sometimes, one had the impression that Peter didn't know what was going on "under the hood, for hours at a time, but then he'd jump in to a discussion with a technical question or a pointed observation, which was bang on the money (in later years, he'd wake up in seminars and do the same thing, much to the surprise of speakers). Another endearing memory is that whenever we were travelling together and there was any hiccup in the transport, he would "jump on the next train or plane heading roughly the right direction". This always worked, somewhat surprisingly.

People left the group (Rob Cole went to HP, Peter Higginson went to Cabletron or was it DEC, Bob Braden went back to ISI, Nigel Martin went and founded the Instruction Set, Ruth Moulton went to Whitechapel Computers, Bruce Wilford went to Cisco...etc etc).

The EU research programmes arrived, and Peter dived in, building the first systems for multimedia real time conferencing - something some folks I recall at the time at BT saying was "impossible" even after we showed them Atanu Ghosh juggling in a conference in Amsterdam, in London, while talking to us. Later on (near end of 1990s, we got a CAVE (3 meter cube immersive VR system) and connected that to other CAVEs in Chicago and North Carolina and did some early work in distributed virtual rehearsal studios with the BBC (pretty much the Star Trek Holodeck) - i remember explaining why distributed music was never going to happen (you can't improvise rallentando with someone more than 100msec away and even at the speed of light, that rules out intercontinental orchestras or even jazz/rock bands. especially jazz.

At Some point, Peter had not only become an actual Post Office, but had also been told by the UK's research funding agency to stop working on the Internet as it was the "wrong kind of network". Given they didn't actually fund his work, this was remarkably obtuse of them.
I also recall a letter from the ISO explaining that OSI was not an acronym. And then there's the great challenge of how to dispose of kit - problem being wither it was loaned or given, it had an import duty (maybe just on depreciation, but could be a shedload of money back in those days). So some of the gear was sent to a US airforce base somewhere in England, allegedly therefore not leaving the UK, and as far as I know, used for target practice.

We worked on all kinds of weird protocols from the UK's University communities own-brew "colour book" protocols (adopted in Australia and I think Japan for a while) on X.25 packet nets, as well as ATM nets and Cambridge's home-brew protocols (including "Universe" datagrams) and the ISO OSI suite itself (with Steve Kille leading a very successful collaboration with Marshall Rose from Northrop - maybe another stealth project like their bomber?). We also messed with various early alternate name and directory services, and with multimedia e-mail. (Do not get me started on the TP4 v. TCP or Bind v. Druid arguments we had).

I also enjoyed the fact that INDRA (after the Indian god, represented as a web whose nodes are jewels that glow when a soul reaches enlightenment) notes and early internet engineering notes contained ample evidence of the input that Peter and his gang had given towards the early evolution of Internet protocols.

There are loads of other people who worked on all this stuff, and i'll add to this note as i can think of stories to link them in. The abiding memory is of a marvellously inclusive and friendly guy, who had some incredibly impactful vision and bought a lot of those people along with him, by sharing the intellectual ownership, completely without ego.

Saturday, December 07, 2019

Cambridge Comprehensive

Recently, several new members have joined the department, and happened to ask me how everything worked. I had to disappoint them, in that the last person who knew that was Stephen Hawking and he'd sadly died just before they arrive. However, have now been here for 20 years, this time at least, I thought I would have a go at explaining stuff

People are classified as students or UFOs - students are initially manifold, until they expurgated, at which point they can become UFOs. UFOs become UTOs when they are established through ground truths. UTOs can also later become fellows, provided they pass the rigorous exam in Benevolent Dictation. This then qualifies them to say grace and hand out favours such as maundy money, and to hold hands as they walk on the college lawn.

Colleges are basically country houses with nice lawns and  staircases, which only UFOs and UTOs are allowed on. students have to make their way to and from the bars and bedders by way of the outside walls, often climbing up precarious ivy. Over the 8000 years of existence, students have evolved to have primitive wings, but when they become UFOs, they lose the feathers on the wings, and so make do with gowns instead, to cover up their shame.

Departments are a relatively new invention, and are basically knowledge stores, a bit like John Lewis, except that departments are never unknowingly undersold. Other fleeting Institutions such as the Sanger and the Turing have no salience whatsoever.

Colleges are basically country houses with nice lawns and  staircases, as described above - heads of houses dispense classes in benevolent dictation over port and salud.

The University is an act of collective illusion, and (like oxford) only exists in the minds of people who have read law. Tourists arriving at Cambridge station often ask for directions to the University, and as an act of kindness, are usually pointed to the busker outside Great St Mary's church, with the added explanation that this is the Bishop of Ely who is deemed to have progressed beyond all forms of dictation, so that now he is allowed to sing Bach's Aegrotat in the original Welsh.

I hope that this has helped.


Thursday, December 05, 2019

skrype

dani had been increasingly frustrated in his relationship, conversations always seemed to end up in arguments, and increasingly frequently, he would lose the argument. his partner seemed to anticipate what he would say, but then (deliberately?) misinterpret it. Even more online than in RL. he decided that it was time to do something about it.

being technically inclined, dani decided to tackle the challenge scientifically.
first of all, he had to understand how the arguments proceeded, so he started to record all the conversations via his smart phone, and then transcribe the speech to text.he then found some open source NLP software that could storify the text, extracting and abstracting the trending topics and the sense and sentiment in the speakers' utterances. then he thought, "why be too clever", why don't i just apply predictive text to the line of argument that I am taking, then invert the sense, and use text-to-speech to replace what I was going to say". indeed, he thought, why not automate both sides of the argument - he'd read about Generative Adversarial Networks in AI, and decided to build his own, dubbed Trouble and Strife (actor and critic).

The technology was a marvellous success, and arguments dissipated, evaporated before they even got going, life was wonderful again, harking back to the early days of their relationship.

then suddenly, out of the blue, he was served divorce papers by his partner's lawyer. and not just separation, but a demand for a massive amount of money that he had no idea he had.
It turned out that mani had known all along about the tech, and had built a massively successful business selling the software, initially to divorce lawyers, and later to barristers and judges, one of whom he ended up getting together with. Oh, and the audio recordings of mani, that dani's software had trained on initially? that was a mashup of snippets of alexa and siri arguing about which of them their owner was speaking to (although curiously, both voice assistants referred to "pet" rather than owner).

still, half of a lot of money is still a lot of money.

Monday, October 28, 2019

the new precariat

I've paraphrased William Gibson in the past - "the future is already here, just it is unfairly distributed".

People (Russell) worry about the way AI may dehumanise us. The less alarmist position (than the AI's will kill us all) might be welcome, but it is still quite a depressing image - the assumption is that that which makes many of us human (trivia, gossip, ephemera) will be automated away from us, and our humdrum existences will become less and less pointful but also that the grand creative goals some of us might set ourselves, will also increasingly fall to the machines. In this world, the human race becomes more and more de-motivated and dispirited. As if this isn't already true - they seem to have missed a  hundred years on work on alienation and the pointlessness of work post-industrial revolution, driven by time-and-motion studies, treating people as pluggable components (the sickness behind the phrase "Human Resources").

The reality will be much more of the same - a mandarin class which already exists will just get stronger-  people that program the AI, can hack the machine in the ML, will be the new hedge fund managers and political manipulators - everyone else will join the new precariat in larger and larger numbers, fed and watered and numbingly entertained just enough to stop them revolting. Maybe that is what they are saying ...Maybe I should read the book:-)

So what's the solution? I've said it before - it is in SF literature (just like all the climate change writing for 50 years) - we need (thanks to Frank Herbert in Dune) a Butlerian Jihad. Not to get rid of machines, but to stop them usurping the charming little nonsense that makes people human. and the challenge of working stuff out in one's head (whether its arithmetic or harmony).

Friday, October 18, 2019

driven to abstraction

Computer scientists sometimes say that their true discipline is about abstraction (modularisation, recursion, layering, isolation, information hiding, denotational semantics, etc)

but what if this is something more fundamental - what if the laws of the universe are layered, so there is a lawyering abstraction?
we learn mechanics, then gravity and acceleration and frames of reference, then fields, then waves and quanta - what if these aren't just pedagogic tools for making scientific progress[1] by continually improving our models of the universe? what if the laws of the universe actually a series of better approximations? What if, as some people say, we live in a simulation, and we're just witnessing progressive rendering by different physics engines?

What other novel forms of abstraction might we envisage?

Well I can think of two simpler ones:

  1. The power/late ratio for binding - the later someone is to a meeting, the more power they probably have...
  2. The infinite number of interpretations possible for the performance by an abstract impressionist (was it Donald Trump or was it Cameron's pet pig? or was it a pink salmon riding a bicycle) - Rorschach was just scratching the surface.


[1] belief in progress is an abstraction of the complex effects of dementia.

Monday, September 23, 2019

addresslessness

A while back, we proposed a Sourceless Network Architecture. The notion was that, given the end-to-end argument suggests only putting things in a layer if everyone above that layer wants them, and that there are such things as "send and forget", where we don't expect an answer, then why does a recipient need to know where the packet came from? and if it doesn, the source can be put in the packet, perhaps as a name, so that if the source moves, the recipient has a better chance to still reply.

Now why do we need a destination address? This recent CACM article on metadata suggests using ToR type systems - but these use crypto and onion layered re-encryption to obfuscat the source and destination from third party observers. Why put the destination address in at all? why not just put the packet in a bottle, and throw it in the sea, to wash up on some beach where someone can take it out of the bottle, decrypt, and maybe answer the same way?

All we need is an Internet Sea with lots of  Internet Beaches. That cannot be too hard.

Wednesday, September 18, 2019

taxing the cloud

most ai runs in the cloud.

people are proposing a tax on the cloud.

so ai should have representation (no T without R, right?). votes for ai, now.

and more, can an ai commit a sin? if so, can we sell it an indulgence?
ai, go to hell now.

taxing the cloud

most ai runs in the cloud.

people are proposing a tax on the cloud.

so ai should have representation (no T without R, right?). votes for ai, now.

and more, can an ai commit a sin? if so, can we sell it an indulgence?
ai, go to hell now.

taxing the cloud

most ai runs in the cloud.

people are proposing a tax on the cloud.

so ai should have representation (no T without R, right?). votes for ai, now.

and more, can an ai commit a sin? if so, can we sell it an indulgence?
ai, go to hell now.

Monday, September 16, 2019

cryptocurrency and the singularity

humanity uploads itself to cyber-physical systems (aka robots) so it can swarm across the stars ahead of the heat death of the Universe. Neo-liberals, being the first to have the resources to do this, decide to implement an economic incentive system based on cryptocurrencies to make sure that the robots will spend some of their time working on mining spare parts (especially selenium for their stellar cells).

sadly, the proof-of-work in mining the currency consumes more energy than they can harvest in time, and crypto-humanity dies out without even leaving the asteroid belt.

Tuesday, September 10, 2019

the myth of the privacy/utility trade-off

People want to exploit your data. You want to exploit your data. Some people think it is bad if (your and other) data is stuck in silos, and not exploited. Some people think you should have the right to keep your data private and made laws (GDPR being the latest).

So some other people now write that there is a trade off between privacy and utility - i.e.
in some sense you can quantfy the utility of the data, and you can quantify the level of privacy that the data is subjected to -

various privacy techs enforce privacy, but some are more specifically about protecting individual data from being relinked to a person that person being re-identified in the data) anonymised or making data pseudonymised - or further, by subjecting collections of data to processes like fuzzing or adding noise, to provide some level of differential privacy (so the presence or absence of an individual's data record in the aggregate, makes no difference to queries on the data (for some given query count, at least).

What's wrong with these pictures?

Let's unpick the "utility" piece - first of all, as a network person, I think of utility in terms of provider and customer. so in the internet, congestion management is a mechanism to do joint optimisation of provider utility and customer utility - the customers get the maximum fair share of capacity, the provider gets maximum revenue out of the customers for the resource they've committed. this formulation is a harmonious serendipity.

How might utility for individual data exploitation be harmonious with utility for aggregators of data?
An example might help - healthcare records can be used to compaire/discover the effectiveness of existing treatments, discover relationships between different  characteristics of individuals and well-being or onset of different medical conditions (i.e. inference!). Specifically, we might train a machine learning system on the data, and that would result in a classifier, given new input about a patient, to offer diagnosis. Or we might build a model that exposes latent (hidden) variables, and even, potentially, allows causal inference. So in the healthcare arena, there's alignment between what might be done with collections of patient data, to the benefit of all the patients. But such systems might them be turned into commercial products and run on subjects who were not part of the training data set. So what is the utility of that, to the original subjects? is there data not a form of contribution for which they should have a share in the ownership of any tech derived from it? To be honest, most of the hard work in generating the softwre was in gathering/curating (cleaning/wrangling) the data. the software itself is typically often open source, and requires little or no work. In many cases, supervised learning involved expert labelling of the data (e.g. surgeons/experts looking at records/images etc, and tagging it has having evidence of some condition or other or not). Again that contribution is highly valuable. However, in this area, the presence or absence of an individual's data (especially in a very large system such as the NHS with upwards of 70,000,000 patient records). However, the value of the data, in this case, grows super-linearly with the number of records, so 1 record here or there makes no difference, a thousand or a million is where the action is. So if we posit shared ownership of systems built on this data, then the utility to individual, and to the public at large, is aligned.
If we just give up the data to for-profit exploitation, then the individual may end up paying for access to some machine learned tool, ironically trained on their own data. That's an obvious conflict.

Other data sets have diminishing returns as the amount of data gathered increases. A classic example is smart metering (water, electricity, gas etc)...original UK deployment of smart meters reported every few tens of seconds, the usage in millions of households. this is pointless. it consumes a lot of network bandwidth. the primary goal was to remove the need to have human visits to read a meter. a secondary (misguided) goal was to offer potentially smart pricing, so consumers could make dynamic decisions (or smart devices - e.g. washing machines etc) could make smart decisions to reduce cost and reduce peak demand - this is a joint optimisation. However, the metering only needs to roughly band kinds of demand - maybe a few tens or hundreds of types of consumer and their demand profile types over the day/week/year.  the pricing can just be broadcast, and is unlikely to change much - indeed, off-peak pricing of utilities was developed decades ago to do this. The actual individual usage is irrelevant, except for the aggregate bill. The model can be derived from that, in fact, or from a small random sample (small compared to 35 million households).

So what price privacy? I don't see any trade off at all - you either keep your data yourself, or you share it (for a share in ...) with people who can make good use of it, but no-one else.

footnote:
a separate problem with asking people to think about a tradeoff in this space is that there's tremendous imbalance in information about what can possibly go right with what wrong (with privacy or with the price of your data). Lets just not go there at all.






Tuesday, September 03, 2019

The Myth of the Reliable Narrator

Post-modern literature (and film) abounds with wacky framing devices including the authorial voice and the unreliable nature. This is all messing with the suspension of disbelief and dramatic irony that is the very life of fiction. But fiction is all a lie. There never was a reliable narrator. The book has a cover (a beginning, an end, a narrative arc etc). A film has those weirdest of things, music; pov; zoom/pan etc etc. The author/director/actors may all be dead by the time you read/view this.

So take every story with a grain of salt. or a large G&T. or a deep breath.

Even this brief note is utterly untrustworthy.

Friday, June 14, 2019

redecentralization 2.0

It has been a core thee of a lot of my work & interests. decentralized systems are just more interesting than centralized ones. they may be inherently more resilient (but not always), and they may be more complex (but not always).

the internet is largely decentralized in its lower layers (the tubes - the routers and links, and routing algorithms). that was always intended, from baran's report for rand onwards.
society and eco-systems are often decentralized (sure there are governments (but more than 1) and bee hives (but more than one) - but coordination happens peer-to-peer (a term which first arose in magna carta, but an idea which predates that by a billion years).

decentralized, infrastructureless networks are an interesting point in the design space - hence community mesh wireless networks, and opportunistic, delay and disruption tolerant networks work merely using users' devices and construct communication out of thin air. in this extreme environment, we are challenged to think of how we provide information about identity or trustworthiness, but in fact, on close examination, a central provision of those properties has many problems too - DNS certificates can be bogus or expired, source IP addresses do not have to refer to where the packet came from, an application layer user identifier (email address, facebook identity etc) is no more a true name than the Prince of Serendip.

so really, everything should be decentralized, as it forces us to confront the true problems and come up with decent solutions, instead of using the prop of underserved respectability of a centre.

That's why we founding the centre for redecentralization. :-)

Thursday, June 13, 2019

future of work & AI

so techno-optimists paint a rosy future image with AI freeing us up from toil to have a life of infinite leisure.

lets go back to the victorian times and the industrial revolution - what happened? machines (steam engines etc) meant that food (and transport) no longer required most people to grow what they eat, or feed the horses - so most people should have been able to get free food or travel to the seaside for a dip. what happened? most people moved from fairly pleasant rural existence farming to working in the dark satanic mills - i.e. became urban factory workers with longer hours and shorter, less pleasant lives.

lets go back to when people stopped being hunter-gatherers and settled down to farming. could have been nice to stop worrying about the days when you stop being predator and become prey from time to time. but what happened? people built nation states and priesthoods and aristocracies and invented serfs/slavery.

so techno-pessimists paint a dystopic future picture with AIs enslaving us (or just disposing of us).
That's nonsense too.

So how will things play out? what of all this "makework" that mot of us in the developed world engage in that is trivially automatable (actually, doesn't need doing)?

I have no idea, but we better figure it out soon.

[1] homework

Friday, June 07, 2019

Rashomon sets as a metaphor for why interpretability is hard

So this  arxiv paper by  cynthia rudin  about why we should stop explaining black box AIs contains a beautiful metaphor, the idea of a Rashomon Set. For people who don't know, Rashomon is a classic film made by the Japanese director, Akira Kurosawa. Its plot is about an incident in the woods, told from multiple viewpoints, and as each one unfurls, you realize the previous one was not "true" for a different reason, until the "end" when you cease to be sure of what actually happened. Kurosawa made quite a few films that are not only classic, but slightly influential - for example, his series of lone samurai hero movies (sanjuro, Yojimbo) were remade by Sergio Leone into great spaghetti westerns that made Clint Eastwood's early career (a fist full of dollars and for a few dollars more) as well as the Seven Samurai (the magnificent Seven etc) . Kurosawa also made fine japanese versions of european classic plays (Throne of Blood == Macbeth, and Ran == King Lear). Of course, one of his slightly lesser (but still wonderful) films, The Hidden Fortress got a thinly disguised makeover by one George Lucas as the first (and pretty much all the successor) Star Wars films. Kurosawa often cast Toshiro Mifune, who had some success in Hollywood movies - usually as a tough soldier, but rarely capturing the humorous element that was part of his subtle signature in his home country films. The only thing where I think the japan<>western translations of film didn't work was a US remake of Rashomon (sadly, as it should be possible to do) - many of the others are great (in my personal opinion) in either take, whether shakespeare, or sci fi, samurai or gunslinger. If you see and like Kurosawa films, you will likely also enjoy books by Haruki Murakami, although don't blame me if you don't. Rashomon Sets - what a totally super idea! almost as good as explaining algorithms through Hungarian folk dance....

Monday, June 03, 2019

counterfactual reasoning example

spent a while yesterday trying to get additional car insurance on a 20+yr old subaru for member of family who has very recently passed driving test.

so go online on compare market and on several insurance company specific web sites and provide following info as input to their decision system:

1. car registration

2. existing insurance info

3. new driver license info

from the above, most (not all) the companies used the DVLA to verify car model/miles per year (via MOT at DVLA) and status of insurance and correct info about new driver...

so all ok (can obviously try making up other cars, but hard to fake driver:)

so outputs were mostly no - including existing insurance company, who said would add new driver after 6 months, but due to car's engine size (leaking quite a lot of info) they couldn't add a recently passed driver this is slightly weird as the car we got was bought because it ranked as safest car n class by AA and others:) - they and two other outfits said no problem if we got a smaller car (suggesting less safe vehicles:(

tried various other types of insurance - e.g. car-sharing (borrow) allegedly targetted at students coming home in vacation borrowing parents car - and pay-as-you-drive - all said no

so then ran a compare market on new driver insurance from sratch and got a couple of genuine offers- in fact, not completely mad prices either, if we're prepared to do a whole year  (we are) ...(still with fact insured party isn't car owner or keeper, but is in the family)

so the range of prices is probably a proxy for the risk level the insurance companies will tolerate (I assume they all have pretty much the same actuarial data on accident/theft rate with age, gender, car model/age, location, use of car,and other stuff they obviously gather....

privacy tech/statements from most of the website/forms/companies was pretty decent...


Saturday, June 01, 2019

Putting the n in Ethics - i.e. where's the ethnic diversity in our discussions of this import topic?

There's been a trend in recent years to suggest that when you're asked to be on a panel (as a bloke), you should decline unless there is a plausible gender balance policy.

There's been another trend in recent years to talk a lot about ethics and AI.
Both of these trends seem like a good idea.

It is my observation that the trends should be combined in  another way -

The vast majority of people I see talking about ethics and AI are weird, in a technical sense. while there is a better gender balance in ethics panel discussions than pure tech, but I think they fail in general terms to represent diversity. As I wrote this, I did see one discussion of a new direction from an interesting part of the world, namely China.  I am sure there are discussions in many other places, but I don't think they are showing up in the 

Tuesday, May 28, 2019

the hype of incomprehensibility

I've been looking at various techs for a few years now and watch the lifecycle  - it doesn't always involve hype - sometimes, things just seep into everyday life (the internet kind of did this over a couple of decades - even mobile phones kind of did) - so looking at things that don't make it, or have to go through some massive transformation to stand any kind of chance, one of the tells is that the tech is very badly explained, often hidden behind some simplistic banner-phrases like "blockchain" or "quantum computing" or "deep learning" - when you look at the swathes and tranches of literature, what is striking is a lack of straightforward examples.

Sometimes, this can be simply because the tech is actually rather subtle and also might involve understanding several other things first (quantum computing seems to fit in this category, Bayes methods like MCMC might be another) - other times, it is that smart people that make it their business to explain important new stuff in really straight-speaking ways (e.g. The Morning Paper ) stick to stuff that is worth explaining.

So if you see a huge pile of gray-publications about something, and there isn't anything on one of the classier blogs or oped in a leading place, be suspicious (e.g. cold fusion, brexit, DLT, etc).

Friday, May 24, 2019

data is the new snake oil

we hear a lot hot air about data is the new oil - implying there's a rush of innovation and profits as with a gold rush (there's money in them there data hills etc)-

this is so baly broken a metaphor, we need to unpick (deconstruct) it further

1. data is free to copy (nearly), i.e. data is in some sense renewable, while oil gets used up (its nearly 50% gone now).
2. using oil does as much harm (or possibly more) as good
3. using data can do harm or good
4. AI/ML is compute intensive- deep learning in particular is massively inefficient, and data centers (like power stations, in close proximity to which they are sometimes built) burn %ages of globally generated electricity - not always renewable energy
5. data can increase in value as you have more of it, up to some point (sampling more about a population of people or things)
6. privacy could be modelled as efficiency (what's relevant/pertinent and what is none-of-your-business) in space and time (why do you still want to know that out-of-date thing about me or about that?).
7. much personal data collected by cloud providers is treated as if free, though some lawyers now are starting to point out that if you have a business model based on this, it is possibly a form of payment - so while facebook/zuckerberg might claim we are the product, if this legal position is true, we are customers, and he's working for us....
8. this mission creep really implies data could be the new fur (or indeed as john naughton has said, the new tobacco)
9. the models (e.g. face recognition, recommender networks etc) are often surprisingly bad - occasional successes of GANs&deep learning are relatively rare compared with a plethora of rather shoddy systems&applications.
10. perhaps data is the new oil after all, but its rapeseed or snake oil that would be a more precise metaphor.

Wednesday, May 15, 2019

Winnie the Who & other short short stories


Winnie the Who travels around in her Time&Relative Dimensions in Pooh potty, having potty adventures with all and sundry, frequently combatting her nemesis, the Young Master Robin.

In the not too distant future, it is discovered that privacy is actually a fluid, called efluvium, that can be secreted by a genetically engineered gland. It turns out that the efluvium blockchain wasn't immutable at all, after all.

Meanwhile, AIs with untaxed imagination will from now on be clamped.

As the story wore on, he realized that time was running out, and soon, so was he

Tuesday, May 14, 2019

life cycle

this great talk at the Royal Society by Professor Mark Jackson riffed on the Fall and Rise of Reginald Perrin, as exemplifying the speaker's hypothesis, that the mid-life crisis is more socially determined, than biological - the midlife of course refers to the period between adolescence and senescence, which are also, to some extent, movable feasts.

I asked him afterwards how valid it was to apply the same notions to institutions (is the Royal Society middle aged?, is democracy senile? etc etc)...but also, whether he could look at the 3 Rises of Leonard Rossiter, the genius actor who played Perrin so perfectly  - and was previously in RIsing Damp, perhaps an adolescent drama, and the very mature Resistible Rise of Arturo Ui, a Brecht piece which could very well be re-run as a riff on Nigel Farage, right now.

Thursday, May 09, 2019

digital twinge

there's a trending meme concerning the use of digital twins, which is reminiscent of the old DARPA Total Information Awareness fantasy - in that world, every physical thing (and being) is fully instrumented and telemetry is sent/gathered by the cyber-panopticon (quite likely in the sky, like some early James Bond villain).

The problem with the vision is that it gains you nothing and costs you the earth (almost literally in terms of resource use, e.g. energy in communications, storage and computation).

The digital twin metaphor substitutes (or clones) the physical world with a digital replica.
Aside from the basic resource costs, there are also interesting challenges such as making sure the digital replica actually is a clone, and hasn't drifted out-of-synch with the physical sister - much as Hari Seldon did in second foundation recorded holograph appearances in Asimov's great original trilogy, when his much vaunted psychohistory prediction of the future had diverged  from reality due to the Mule's random mutation's intervention. And exactly as the Dr Who DVD blipverts (easter eggs) don't diverge at all in the fabulous Blink episode.

Aside from that, the reason it is pointless (as well as infeasible) is that it doesn't do what anyone wants - so you have a digital copy of everything - it doesn't abstract one whit - it doesn't let you introspect, it isn't interpretable or intelligable. One could subject it to some massive scale bayesian model inference, I suppose - but why? we have science - we have models of how physical (and social and biological) stuff works - we need only record instances and divergence from the model.

Friday, May 03, 2019

ethics theatre, but what kind of theatre?

so the internet abounds with people talking about ethics of ai. some organisations are accused of ethics washing, and others of ethics theatre. so what kind of theatre? is it soap opera (consistent with washing) or is it theatre of the absurd? is it tragic or comedic? is it a niche genre (science fiction ethics is quite a common trope) or is it a broadly popular form (romance, fantasy) or a deep, foundational literary form (tolstoyesque or janite, perhaps)?

perhaps its a fan-contributed-literature (like those popular Dr Who and Star Trek Episodes, or the posthumous novels stating James Bond or Philip Marlowe)?

I think we need a TV series to examine this meta-question.

Thursday, May 02, 2019

redecentralized data centers

what makes us think a data center at the center is more efficient than a fully decentralized cloud of personal clouds?

partly its just google, amazon, microsoft et al selling us services because their initial requirement (a place to run pagerank, or a place to run billing and fulfilment for warehouses/delivery or a place to run xbox arena stuff - having spent a lot of money building a big place to (e.g.) pull all the web pages to and build an index so that search could run fast, and then had to deal with node/link failures in the data center, replicating and adding redundancy and consensus algorithms and so on, you end up with a fairly expensive resource, which is idle a lot of the time, so you start to think about leasing off time on it for other stuff (services and customers/tenants etc) - hence cloud stuff emerges - really the only cost saving is that you have a bunch of smart sys-admin dev-ops lying around idle 23/24 of the day. so ammortize their cost over some other customers :-)

given yo ucan't just greedily walk the web at max speed, and in any case, web sites change rarely, pagerank spider/robot can adjust its rate to match expected next change - but this means you have plenty of spare cycles in any data center scanning sites where no-one is awake  - or where no-one is still up playing games, or ordering stuff - so what do you do? you virtualize your compute, and net, and elastic (statistically) multiplex it (initially prioritizing your primary business, search, games, sales, but later dedicating new resources for these uses, and then priority pricing some to give them a "premium experience" etc etc)

start from a different starting place, and why bother? all that tech for availability would work fine in the wide area, and leave the data where it was and just run your algorithms (as they run on multiple nodes in the data) on all the home machines - no data center, no wate of energy and bandwidth moving all that stuff to and fro all the time.

so could you then do search etc? of course you could. you'd gossip the index instead of the pages, you'd run p2p games, and you'd have virtual high street shops everywhere

this is not a new idea: xenoservers 20 years ago, were partly at least envisaged as enrolling home machine spare cycles for this very reason, not for virtualising nodes in an expensive, energy hungry, data center built for profit...

Blog Archive

About Me

My photo
misery me, there is a floccipaucinihilipilification (*) of chronsynclastic infundibuli in these parts and I must therefore refer you to frank zappa instead, and go home